Research

Grok 4.7 on Senior SWE-Bench: Strong pass@3, Cheaper Cost per Trial

September 21, 2026
3 min read
Snorkel Team

Grok 4.7 was evaluated on Senior SWE-Bench, our benchmark for measuring whether coding agents work like senior engineers. This model improves on Grok 4.6 on both tasteful measures, ranks fifth on tasteful pass@3, and does it at roughly 1/14th of the cost.

[Line Graph of Model Performance]

Benchmark

Senior SWE-Bench evaluates agents as compared to a senior engineer. Instructions are under-specified to be more like a message from a colleague rather than a requirements document. Bug tasks are sourced from real PRs that needed actual runtime investigation to resolve. And scoring rewards taste, not just correctness. A tasteful solve requires verifiers and validation to pass, a quality rubric above 0.5, bloat under 2×, and adherence to load-bearing codebase practices not necessarily present in the task instructions.

Tasks are drawn from real PRs across 50 public and 50 private problems spanning different libraries and multi-service applications. The benchmark is hard by construction: the strongest frontier models fail to produce a tasteful solve on more than 65% of tasks.

All models run under Mini-SWE-Agent. Results are reported as tasteful pass@1 and pass@3, alongside average output token cost per trial. Grok 4.7 results below are at xhigh effort.

Results

Grok 4.7 improves on Grok 4.6 on both the tasteful pass@1 and pass@3. Tasteful pass@1 rises from 26.3% to 27.4% and tasteful pass@3 rises from 38.9% to 40.0%. The pass@3 result places it fifth overall, tied with Opus 4.8.

The cost is notably lower compared to other leading models. Grok 4.7 and Opus 4.8 post identical tasteful pass@3 rates, but Grok 4.7 reaches that at $0.24 per trial against $3.36, a 1/14th of the cost per trial. Compared to GPT-5.6 Sol, Grok 4.7 matches its pass@3 rate at 63.2% and is only three points behind on its tasteful pass@3. Grok 4.7 does this though at roughly a quarter the cost per trial.

The plain pass rates move in the other direction, with pass@1 at 49.5% against Grok 4.6’s 51.6% and pass@3 at 63.2% against 65.3%. Grok 4.7 solves slightly fewer problems than its predecessor while producing more solutions that clear the taste gates.

The pass@1 picture is a little bit different. Grok 4.7 achieves a tasteful pass@1 of 27.4%. This is a 1.1-point improvement over Grok 4.6, but it sits below Opus 4.8 at 30.5% and GPT-5.5 at 29.5%, both of which Grok 4.7 matches or beats at pass@3.

Continuing, Grok 4.7 reaches a tasteful pass@1 of 27.4% against a tasteful pass@3 of 40.0%. GPT-5.6 Sol reaches 34.7% and 43.2%, and Opus 5 reaches 34.7% and 45.3%. This shows that the model can perform well across multiple runs, but that its consistency is lower than some of the other model families. With this, Grok 4.7 has the capability to solve the benchmark tasks but landing it on the first shot may be difficult.

This is broadly in line with what Musk has stated about the model before its release. It can abandon hard tasks early and has challenges rigorously checking its own work. Senior SWE-Bench is built for long-horizon work, and the shape of these results suggests that reinforcement learning could be further improved.

Conclusions

Senior-level engineering remains an open frontier as even the most performant models still struggle on many of the tasks. With that in mind, Grok 4.7 is a true improvement over 4.6. It sits just outside the leader group on the benchmark’s headline capability measure while costing between a quarter and a 1/14th as much per trial as every model ranked above it. While it still has limitations on first attempts, the performance is an improvement, and at a fraction of the cost. For teams running agents in loops, where retries are not a major concern, that combination may matter more than first-attempt performance alone.

You can see the full results on the leaderboard at Senior SWE-Bench.

Share this article

Recommended articles

View all articles
Image
Building the Frontier Lab for Agentic Data
Introduction Today, I’m incredibly excited to announce Snorkel’s $[350]M Series E financing [and employee tender] at a $3.5B valuation, co-led by Insight Partners and S32, with significant participation from existing investor Addition. The round included new investors March Capital, Blumberg Capital, Allegis Capital, Frontline, Standard, One Prime Capital, Third Point Ventures, and D.E. Shaw Ventures, along with existing investors Greylock,
September 20, 2026
Alex Ratner
Image
From Foundational Competency to Expert Performance: A Curriculum Approach to Model Development
A student does not go from 1st to 12th grade in a single step. Each grade builds on a specific set of skills, and each one assumes the previous skills have already been mastered. Nobody learns calculus without algebra. When a student skips ahead anyway, what they end up with is memorization rather than understanding. The gaps show up later,
September 15, 2026
Snorkel Team
Image
Terminal-Bench 4.0: Why Continuous Benchmarks Require Continuous QA
The speed of new frontier model releases keeps accelerating. Meanwhile benchmarks struggle to keep up and saturate quickly, often being left in the dust. Most benchmarks are static datasets with no active maintenance, causing them to lose value fast. Some benchmarks are looking to change this by becoming Continuous Benchmarks. Terminal-Bench is one of the most widely reported benchmarks on
August 27, 2026
Justin Bauer
Image
Image

Join our newsletter

For expert advice, the latest research, and exclusive events.
By submitting this form, I acknowledge I will receive email updates from Snorkel AI, and I agree to the Terms of Use and acknowledge that my information will be used in accordance with the Privacy Policy.