Stanford’s Human-Centered AI Institute released its annual AI Index report this week — and the numbers pull in opposite directions at the same time. AI can now solve real-world coding problems at near-perfect accuracy that would have been impossible 12 months ago. Global investment hit $581 billion in 2025. Adoption is moving faster than the internet ever did.
And yet: only 23% of Americans believe AI will be good for their jobs. The transparency scores of the major AI labs dropped sharply in a single year. Documented AI harms rose 55%. And the US-China performance gap that once felt like a structural advantage has narrowed to 2.7%.
The 2026 Stanford AI Index the most comprehensive annual measure of where AI actually stands — captures a field outrunning everything built to manage it.
The Capability Numbers Are Historic
The clearest signal of how fast AI is moving: one year ago, the best models scored around 60% on SWE-bench Verified, a benchmark testing whether AI can autonomously fix real bugs in production codebases. As of early 2026, that number is near 100%.
On Humanity’s Last Exam — a graduate-level test spanning physics, chemistry, law, and philosophy, specifically designed to be beyond AI’s reach — top models from Anthropic and Google now score above 50%. A year ago, the best score was under 10%.
But the report also documents what researchers call the “jagged frontier.” The same models that win gold at the International Mathematical Olympiad read analog clocks correctly just 50.1% of the time. AI agents leapt from 12% to 66% task success on OSWorld, which tests agents on real computer tasks across operating systems — a massive jump, but still a third of the tasks remain out of reach.
The inconsistency matters. Deployment decisions that don’t map to where the frontier is jagged are where the failures happen.
The Trust Gap Between Experts and Everyone Else Is Now 50 Points
Here’s the figure that should concern policymakers more than any benchmark: 73% of AI experts believe AI will have a positive impact on jobs. Among the American public, that number is 23%.
A 50-point confidence gap between the people building the technology and the people living with its consequences isn’t a messaging problem. It’s a structural one.
The US also ranks last among surveyed countries in public trust in its government to regulate AI — at 31%. EU regulatory institutions are trusted significantly more, even among American respondents.
| Sentiment Measure | AI Experts | US General Public |
|---|---|---|
| AI will positively impact jobs | 73% | 23% |
| Trust government to regulate AI | — | 31% (last globally) |
| Excited more than concerned about AI | 56% | 10% |
Source: Stanford HAI, 2026 AI Index — Public Opinion
The employment data gives the public’s skepticism a grounding. Jobs held by software developers aged 22–25 have dropped nearly 20% since 2024. Customer service — another high-AI-exposure category — shows the same pattern. Senior and mid-career positions have held steadier, but the entry-level pipeline is contracting.
The US-China AI Race Is Now Separated by 2.7%
The United States leads in private AI investment — $285.9 billion in 2025, compared to China’s $12.4 billion. It still produces more top-tier models: 50 notable AI systems in 2025 versus China’s 30.
On raw model performance, the lead is 2.7% Anthropic’s leading model holding a thin edge as of March 2026. US and Chinese models have traded the top benchmark position multiple times since early 2025. DeepSeek-R1 briefly matched the best US model in February 2025.
One number that doesn’t make the investment headlines: the US has lost roughly 89% of its incoming AI researchers since 2017, with an 80% drop in the last year alone. Investment dominance and talent pipeline health are different metrics pointing in different directions.
AI’s Environmental Price Tag Is a State’s Worth of Power
AI data center power capacity globally has reached 29.6 gigawatts — roughly what New York State draws at peak demand. That’s current installed capacity, not a projection.
Training a single frontier model is now an environmental event. Grok 4’s training run produced an estimated 72,816 tons of CO2 equivalent — the same as driving 17,000 cars for a full year.
Water consumption is climbing alongside energy. Cooling and hydroelectric demands from GPT-4o inference alone may exceed the drinking water needs of 12 million people annually. The US hosts 5,427 data centers — more than 10 times any other country — and consumes more AI-related energy than any nation on earth.
These numbers rarely appear alongside capability benchmarks in the same breath. They should.
Safety Is Slipping. Transparency Is Getting Worse.
Documented AI incidents rose from 233 in 2024 to 362 in 2025 — a 55% increase in a single year. The OECD’s AI monitoring system recorded a peak of 435 monthly incidents in January 2026.
The organizations handling those incidents are getting worse at it. The share of companies rating their AI incident response as “excellent” dropped from 28% in 2024 to 18% in 2025.
The transparency picture is equally concerning. The Foundation Model Transparency Index — which grades how openly labs disclose training data, compute usage, capabilities, and risks — saw average scores fall from 58 to 40 in a single year. Frontier labs are disclosing less, not more, as the technology becomes more consequential. On jailbreak testing, safety performance dropped across all major models evaluated.
Frequently Asked Questions
What is the Stanford AI Index 2026?
The Stanford AI Index is an annual report produced by Stanford University’s Institute for Human-Centered AI (HAI). Now in its seventh edition, it tracks AI technical performance, investment flows, adoption rates, public opinion, workforce impacts, and safety trends across the global landscape. The 2026 report was released in April 2026 and is freely available at hai.stanford.edu.
How fast are AI capabilities improving in 2026?
Significantly. Performance on the SWE-bench coding benchmark rose from 60% to near 100% in one year. Top models now score above 50% on Humanity’s Last Exam, up from under 10% a year ago. AI agents also jumped from 12% to 66% task success on real-computer-use benchmarks. The same models, however, remain inconsistent on simpler perceptual tasks.
Is China catching up to the US in AI?
Yes — meaningfully. The performance gap between the best US and Chinese AI models stands at 2.7% as of March 2026. The two countries have traded the top benchmark position multiple times in the past 12 months. The US leads in private investment by a factor of 23, but China leads in research publication volume, patent output, and industrial robot installations.
What does the Stanford AI Index 2026 say about AI and jobs?
The picture is split. Measurable productivity gains exist 26% improvement in software development output, 14–15% in customer support. But employment for entry-level software developers aged 22–25 has dropped nearly 20% since 2024. Only 23% of Americans believe AI will benefit their job situation, compared to 73% of AI experts.
Where This Is Heading
The Stanford AI Index doesn’t editorialize. It measures. And what it’s measuring in 2026 is a technology moving faster than the institutions, regulations, and public confidence built to accompany it.
The capability gains are real. So are the safety gaps, the transparency declines, the environmental costs, and the trust collapse outside expert circles. What the report makes clear — and what both policymakers and company boards should take seriously — is that these aren’t separate stories. They’re the same story, measured from different angles.
The AI field has spent years arguing that speed and safety aren’t in tension. The 2026 data makes that argument harder to sustain.