AI Coding Benchmarks Are Broken — Here’s What Stanford Found

· By AIX Cove · Reviewed by AIX Cove · ai-coding-development
AI Coding Benchmarks Are Broken — Here’s What Stanford Found

AI Models Now Solve Nearly Every Coding Benchmark, and Stanford’s Latest Report Shows What That Actually Means

Something odd happened in AI coding benchmarks over the past year. Performance on SWE-bench Verified, the standard test for autonomous software engineering, went from roughly 60 percent to near 100 percent. One year. That kind of jump doesn’t happen in mature fields. It barely happens in anything.

This datapoint comes from the 2026 AI Index Report published by Stanford’s Human-Centered Artificial Intelligence center earlier this week. The ninth edition of the annual report spans over 400 pages and covers everything from compute capacity to carbon emissions to public opinion. But the coding benchmark numbers stand out, because they suggest the gap between “AI can help write code” and “AI can write production code” has basically closed.

Ray Perrault, co-director of the AI Index steering committee, urges caution when interpreting benchmark results. “We generally lack measures of how well a system (or agent) needs to function in a particular setting,” Perrault told IEEE Spectrum. “Knowing that a benchmark for legal reasoning has 75 percent accuracy tells us little about how well it would fit in a law practice’s activities.”

He’s right. A benchmark is a controlled environment. Production codebases are not. But the speed of improvement on SWE-bench Verified is still worth paying attention to, because it tracks something closer to real work than, say, multiple-choice trivia. The test asks models to resolve actual GitHub issues from popular open-source repositories. Going from 60 to near-perfect in twelve months means the models got substantially better at understanding messy, human-written code and figuring out what went wrong.

Humanity’s Last Exam Isn’t Lasting Very Long

The coding benchmark isn’t the only one getting crushed. Humanity’s Last Exam, a benchmark built from expert-submitted questions designed to represent the hardest problems across academic fields, tells a similar story. In the 2025 AI Index, the top-scoring model (OpenAI’s o1) answered just 8.8 percent of questions correctly. The 2026 report puts that figure at 38.3 percent. And as of April 2026, models like Anthropic’s Claude Opus 4.6 and Google’s Gemini 3.1 Pro have already crossed the 50 percent threshold.

At this rate, the exam’s name is aging poorly. The trend is clear: whenever researchers build a benchmark meant to “finally” challenge AI models, the models catch up faster than anyone expected.

Who’s Building All These Models

The United States still leads in raw model output. Epoch AI tracked 50 “notable” model releases from US-based organizations in 2025. But the gap with China has narrowed to the point where the two countries trade the lead on specific benchmarks multiple times within a single year. The 2026 report explicitly notes that the US-China model gap has effectively closed.

One trend that hasn’t changed: almost everything comes from industry now. Epoch AI logged 87 notable models from companies in 2025, compared to just seven from academia and government combined. That’s over 90 percent. In 2015, the split was roughly even. In 2003, it was zero industry models. The pipeline from university labs to corporate product teams is now running in one direction.

The Compute Build-Out Behind the Numbers

None of this progress is free. The report includes EpochAI’s estimate of global AI compute capacity, measured in H100e-equivalent units. The total has grown more than threefold every year since 2022. Since 2021, that’s a 30x increase.

Nvidia sits on top of this pile, with its GPUs accounting for over 60 percent of total AI compute capacity worldwide. Amazon and Google, both designing their own AI chips, occupy second and third place. TSMC’s upcoming earnings results on Thursday are expected to confirm the $1.7 trillion company’s continued dominance in manufacturing the chips powering this expansion, though Japan’s $16 billion bet on Rapidus and an Intel partnership with Elon Musk suggest not everyone is comfortable with Taiwan’s grip on advanced chipmaking.

The Carbon Problem Nobody Wants to Quantify Precisely

Here’s the uncomfortable part. Training a frontier model like xAI’s Grok 4 generates an estimated 72,000 tons of carbon-equivalent emissions, according to the Stanford report. For comparison, GPT-4 was estimated at 5,184 tons, and Meta’s Llama 3.1 405B came in at 8,930 tons. The trajectory is not subtle.

Perrault notes these figures carry significant uncertainty. “These estimates should be interpreted with caution. In the case of Grok, they rely heavily on inferred inputs drawn from public reporting, xAI statements, and other non-verifiable sources,” he says. But he also points out that Epoch AI independently estimated Grok 4’s emissions at roughly 140,000 tons of CO2, which is nearly double the report’s conservative figure.

Inference energy use varies wildly between models. DeepSeek’s V3 consumes around 23 watts per medium-length prompt, while Claude 4 Opus uses about 5 watts. The least efficient models produce over ten times the carbon per query compared to the most efficient ones. As deployment scales up (88 percent organizational adoption, per the report), the inference side of the carbon equation could end up dwarfing training costs.

What 88 Percent Adoption Actually Looks Like

The report found that 88 percent of organizations surveyed have adopted AI tools in some form, and four out of five university students now use generative AI regularly. Those are adoption numbers that most technologies take decades to reach.

China, meanwhile, leads the world in robotics deployment by a wide margin. The country installed 295,000 industrial robots in 2024, compared to roughly 44,500 in Japan and 34,200 in the United States. The AI Index has always tracked both software and hardware, and the robotics numbers are a useful reminder that “AI adoption” means very different things depending on which country and which industry you’re looking at.

The US has another emerging problem: talent retention. The largest share of identified AI authors and inventors came from the United States in 2025 (220,520 people), followed by India (50,460) and Germany (48,520). But the report notes that the US is “finding it harder to attract top talent” even as it outspends every other country on AI research and development. India, despite having a 50,000-strong AI talent pool, leads the world in net talent outflows.

So What

The 2026 AI Index confirms what anyone following the field already sensed: model capabilities are accelerating, adoption is near-universal in enterprise settings, and the environmental costs are scaling just as fast as the benchmarks. The gap between US and Chinese model performance is gone. Coding benchmarks are basically maxed out. Humanity’s Last Exam is halfway beaten.

The real question isn’t whether AI can pass tests. It’s whether the infrastructure, the energy grid, and the regulatory frameworks can keep up with what the models are now capable of doing in production. Based on this report, the models are ahead.

For tool-by-tool comparisons, see our AI coding listings and the comparisons section.

Sources: official docs & pricing pages, hands-on testing where noted, and community feedback. Prices verified August 2026 and may change.