Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. This model improves on Grok 4.6 on both tasteful measures, ranks fifth on tasteful pass@3, and does it at roughly 1/14th of the cost.
[Line Graph of Model Performance]
Benchmark
Senior SWE-Bench evaluates agents as compared to a senior engineer. Instructions are under-specified to be more like a message from a colleague rather than a requirements document. Bug tasks are sourced from real PRs that needed actual runtime investigation to resolve. And scoring rewards taste, not just correctness. A tasteful solve requires verifiers and validation to pass, a quality rubric above 0.5, bloat under 2×, and adherence to load-bearing codebase practices not necessarily present in the task instructions.
Tasks are drawn from real PRs across 50 public and 50 private problems spanning different libraries and multi-service applications. The benchmark is hard by construction: the strongest frontier models fail to produce a tasteful solve on more than 65% of tasks.
All models run under Mini-SWE-Agent. Results are reported as tasteful pass@1 and pass@3, alongside average output token cost per trial. Grok 4.7 results below are at xhigh effort.
Results
Grok 4.7 improves on Grok 4.6 on both the tasteful pass@1 and pass@3. Tasteful pass@1 rises from 26.3% to 27.4% and tasteful pass@3 rises from 38.9% to 40.0%. The pass@3 result places it fifth overall, tied with Opus 4.8.


The cost is notably lower compared to other leading models. Grok 4.7 and Opus 4.8 post identical tasteful pass@3 rates, but Grok 4.7 reaches that at $0.24 per trial against $3.36, a 1/14th of the cost per trial. Compared to GPT-5.6 Sol, Grok 4.7 matches its pass@3 rate at 63.2% and is only three points behind on its tasteful pass@3. Grok 4.7 does this though at roughly a quarter the cost per trial.
The plain pass rates move in the other direction, with pass@1 at 49.5% against Grok 4.6’s 51.6% and pass@3 at 63.2% against 65.3%. Grok 4.7 solves slightly fewer problems than its predecessor while producing more solutions that clear the taste gates.
The pass@1 picture is a little bit different. Grok 4.7 achieves a tasteful pass@1 of 27.4%. This is a 1.1-point improvement over Grok 4.6, but it sits below Opus 4.8 at 30.5% and GPT-5.5 at 29.5%, both of which Grok 4.7 matches or beats at pass@3.
Continuing, Grok 4.7 reaches a tasteful pass@1 of 27.4% against a tasteful pass@3 of 40.0%. GPT-5.6 Sol reaches 34.7% and 43.2%, and Opus 5 reaches 34.7% and 45.3%. This shows that the model can perform well across multiple runs, but that its consistency is lower than some of the other model families. With this, Grok 4.7 has the capability to solve the benchmark tasks but landing it on the first shot may be difficult.


This is broadly in line with what Musk has stated about the model before its release. It can abandon hard tasks early and has challenges rigorously checking its own work. Senior SWE-Bench is built for long-horizon work, and the shape of these results suggests that reinforcement learning could be further improved.
Conclusions
Senior-level engineering remains an open frontier as even the most performant models still struggle on many of the tasks. With that in mind, Grok 4.7 is a true improvement over 4.6. It sits just outside the leader group on the benchmark’s headline capability measure while costing between a quarter and a 1/14th as much per trial as every model ranked above it. While it still has limitations on first attempts, the performance is an improvement, and at a fraction of the cost. For teams running agents in loops, where retries are not a major concern, that combination may matter more than first-attempt performance alone.
You can see the full results on the leaderboard at Senior SWE-Bench.
Recommended articles







