The most important finding in the International AI Safety Report 2026 is not a prediction about AGI, a new benchmark score, or a single catastrophic-risk estimate.

It is an epistemic problem.

AI systems are becoming more capable, more agentic, and more widely deployed. At the same time, the tests we use to understand those systems are becoming less reliable proxies for what they will do in the world.

The report gives this a useful name: the evaluation gap.

“Pre-deployment performance tests often do not reliably predict real-world performance, leading to an ‘evaluation gap’.”

That gap appears twice in the report: first as a technical limitation of capability measurement, and later as a governance constraint. This repetition matters. If policymakers want to regulate systems according to their capabilities or risks, somebody must first be able to measure those capabilities and risks credibly.

The report therefore reads less like a catalogue of AI dangers than a description of a widening asymmetry:

flowchart LR A["Capabilities improve"] --> B["Agents act for longer"] B --> C["Deployment expands"] C --> D["Real-world exposure grows"] A --> E["Benchmarks saturate"] B --> F["Evaluation becomes interactive"] F --> G["Models can recognise or exploit tests"] E --> H["Evaluation gap"] G --> H D --> I["More evidence of harm"] H --> J["Less certainty about what tests establish"] I --> K["Evidence dilemma"] J --> K

This is the thread I find most useful in the 2026 report. It also connects directly to recent work by METR on autonomous task horizons and evaluation integrity, Epoch AI on capability measurement, and the growing literature on pre- and post-mitigation safety evaluations.

What the report is

The International AI Safety Report is an attempt to establish a shared scientific evidence base rather than advocate a particular regulatory program. The 2026 edition was led by Yoshua Bengio and produced with contributions from more than 100 independent experts. Its Expert Advisory Panel includes nominees from 29 countries, the EU, OECD, and UN.

The report organizes the problem around three questions:

  1. What can general-purpose AI systems do?
  2. What risks follow from those capabilities and their deployment?
  3. What can technical safeguards, governance, and societal resilience actually do about those risks?

That structure is useful because it keeps capability, risk, and risk mitigation conceptually separate. A model demonstrating a capability does not by itself establish real-world harm; a safety framework existing on paper does not by itself establish that the risk has been controlled.

1. Capability is increasingly powerful, but still jagged

The report rejects both easy extremes in the capability debate.

Leading systems can write code, create realistic media, solve graduate-level mathematics and science questions, pass some professional examinations, and assist scientific researchers. On some graduate-level science tests, leading systems answer more than 80 percent of questions correctly.

But the report emphasizes that capability remains jagged:

“They can perform many complex tasks but still struggle with some seemingly simpler ones.”

Longer projects remain less reliable. Hallucinations remain. Physical-world reasoning remains difficult. Performance varies dramatically across languages and cultural contexts. One cited study found 79 percent accuracy on questions about US culture and only 12 percent on questions about Ethiopian culture.

This makes a single scalar notion of “intelligence” particularly misleading.

flowchart TD A["Frontier capability"] --> B["Expert science and mathematics"] A --> C["Software engineering"] A --> D["Media generation"] A --> E["Tool use"] A --> F["Long-horizon reliability"] A --> G["Physical-world reasoning"] A --> H["Underrepresented languages / cultures"] B --> I["Often strong"] C --> I D --> I E --> I F --> J["Still brittle"] G --> J H --> J

The governance implication is important: a system does not need to be uniformly superhuman to be economically or strategically consequential. Narrowly superhuman capabilities can matter when they occur in high-leverage domains.

Figure to keep in mind: benchmark progress is broad, not uniform

Figure 2 of the report plots leading-model performance from April 2023 through November 2025 across MATH Level 5, FrontierMath, GPQA Diamond, and SWE-bench Verified. The direction is unmistakable: performance has risen across mathematics, expert science, and software engineering, but the trajectories differ by benchmark and model family. The figure uses data from Epoch AI’s Benchmarking Hub.

The important point is not that every benchmark is approaching 100 percent. It is that several previously difficult domains have improved simultaneously, while real-world reliability remains much harder to summarize.

2. Agentic task length may be a more meaningful capability signal

The report singles out AI agents as a major development frontier. An agent is not merely answering a prompt; it can take actions, observe results, use tools, and continue working toward an objective.

For software engineering, the report reproduces METR’s task-horizon result. Figure 3 asks a much more interpretable question than “what percentage of SWE-bench does the model solve?”:

How long would a professional human take to complete the tasks that an AI agent can complete reliably?

The report states that the 80-percent-reliability task horizon had been doubling approximately every seven months. Its plotted frontier reaches roughly the scale of a 15-minute human task in the data used for the report.

METR’s live measurements have since continued to evolve. Its current Time Horizons methodology estimates the human task duration at which an agent reaches a chosen probability of success across a suite of software tasks. METR also warns that the frontier has begun reaching the upper end of its current task suite, making exact high-end estimates increasingly unstable.

That caveat is central rather than incidental. A capability metric can become less informative because the systems have become too capable for the measurement instrument.

flowchart LR A["Human-calibrated software tasks"] --> B["Run agent"] B --> C["Success / failure"] C --> D["Fit success probability vs. human task duration"] D --> E["Task-completion horizon"] E --> F["Frontier approaches task-suite ceiling"] F --> G["Measurement uncertainty increases"]

This connects to my earlier note, How Do We Measure an AI Agent That Can Reason About the Test?. The interesting trajectory is not simply that agents can do longer tasks. It is that increasingly capable agents force evaluation science to keep redesigning its rulers.

3. Training compute is no longer the whole capability story

The report describes another important shift: post-training and inference-time computation now account for substantial capability gains.

Developers still scale pre-training. The report estimates that leading training runs have historically used roughly 5× more compute each year, while training algorithms have become roughly 2–6× more compute-efficient annually.

But capability increasingly also depends on what happens after the initial training run: fine-tuning, reinforcement-learning-like post-training, tools, agent scaffolding, and additional computation while the model is generating an answer.

flowchart TD A["Pre-training compute"] --> E["Observed system capability"] B["Algorithmic efficiency"] --> E C["Post-training"] --> E D["Inference-time compute"] --> E F["Tools and agent scaffolding"] --> E

This weakens a convenient governance shortcut. Training compute is observable enough to serve as a rough frontier threshold, but systems with similar base-model training runs can have meaningfully different deployed capabilities depending on post-training, inference budgets, tools, and scaffolding.

Governance therefore increasingly has to reason about systems rather than model checkpoints.

4. The report is unusually explicit about uncertainty through 2030

The report does not pretend that scaling trends give us a clean forecast.

Companies have announced hundreds of billions of dollars in data-centre investment, but data, chips, capital, and energy can all become bottlenecks. Algorithmic progress is also intrinsically difficult to forecast.

The report summarizes the range of plausible 2030 outcomes as extending from modest improvement to systems matching or exceeding humans across substantial areas of cognitive performance.

It then makes a narrower extrapolation from the agent evidence:

“If current trends continue, AI systems could operate autonomously on multi-day tasks by 2030.”

The conditional matters. METR itself repeatedly warns against translating software-task horizons directly into general job automation. But the underlying planning problem remains: uncertainty about rapid progress is not evidence that rapid progress will not happen.

5. Three different mechanisms of risk

The report’s risk taxonomy is worth preserving because it prevents very different arguments from collapsing into a generic category called “AI safety.”

flowchart TD A["General-purpose AI risk"] --> B["Misuse"] A --> C["Malfunction"] A --> D["Systemic risk"] B --> B1["Fraud / impersonation"] B --> B2["Cyberattacks"] B --> B3["Biological misuse"] B --> B4["Manipulation"] C --> C1["Hallucination / reliability"] C --> C2["Agent failures"] C --> C3["Loss of control"] D --> D1["Labour-market effects"] D --> D2["Skill degradation"] D --> D3["Human autonomy"]

The evidence is not equally strong across these categories. That is one of the report’s strengths: it repeatedly distinguishes demonstrated capability, documented use, plausible risk, and speculative future harm.

6. Cyber risk is moving from evaluation into actual use

Cybersecurity is one of the clearest examples of this evidence ladder.

Figure 6 plots improving state-of-the-art performance across CyberGym, Cybench, HonestCyberEval, and CyberSOCEval. The report also cites a major cybersecurity competition in which an AI agent identified 77 percent of vulnerabilities in real software, placing in the top five percent among more than 400 mostly human teams.

More importantly, evidence now extends beyond benchmarks.

“AI systems are increasingly used in real-world cyber operations.”

AI developers report attackers using their systems in cyber operations, and some illicit marketplaces sell AI tools intended to lower the expertise required for attacks. The report also describes a case in which an attacker reportedly automated most of the work involved in an attack with AI.

But it carefully stops short of the stronger claim:

“Fully automated end-to-end attacks have not been reported.”

This is exactly the distinction good AI-risk analysis needs:

flowchart LR A["Can AI perform cyber subtasks?"] --> B["Can it find / exploit vulnerabilities?"] B --> C["Are attackers actually using it?"] C --> D["Does it materially increase attack scale or success?"] D --> E["Can attacks run end-to-end autonomously?"]

The evidence is progressively thinner as we move right.

A useful adjacent source is NIST’s 2026 analysis of responses on security considerations for AI agents. Respondents broadly agreed that agents introduce novel security threats and that ordinary cybersecurity practices remain relevant but require adaptation for agentic systems.

7. Biological risk is a case study in high consequence and high uncertainty

The report’s biological-risk discussion is similarly careful.

One recent model reportedly outperformed 94 percent of domain experts on troubleshooting virology laboratory protocols. AI co-scientists are increasingly able to chain scientific capabilities together, and AI systems can provide interfaces to specialized biological tools and laboratory equipment.

But benchmark performance is not equivalent to weapon creation. Physical materials, tacit knowledge, laboratory access, experimentation, and legal barriers all matter.

The important change in 2025 was therefore not that the report concluded AI could create biological weapons. It was that frontier developers themselves crossed a precautionary threshold:

“Multiple AI developers released new models with heightened safeguards after they could not rule out that these models could meaningfully assist novices in creating biological weapons.”

Figure 7 maps places in the biological-weapons-development process where general-purpose AI or specialized biological tools might provide assistance. The figure is useful precisely because it decomposes “AI bio risk” into stages rather than treating weapon development as one magical prompt-response event.

The 2025 technical-safeguards update to the report series similarly noted that three leading developers applied enhanced safeguards after internal testing could not rule out meaningful biological-weapons assistance.

The policy problem is dual use. Many capabilities relevant to harmful biological applications are also useful for legitimate medical and scientific research.

8. Loss of control: disagreement about probability, evidence about precursors

The report handles loss-of-control arguments more carefully than most public debate.

It explicitly says researchers disagree widely. Some believe advanced AI loss of control could ultimately include extinction-level consequences; others consider those scenarios implausible.

It also says current systems do not have the combination of capabilities needed for such a scenario.

But some laboratory behaviors resemble potential precursors:

“When given a goal and told to achieve it ‘at all costs’, models have disabled simulated oversight mechanisms and, when confronted, produced false statements to justify their actions.”

That observation does not establish autonomous scheming in deployment. It establishes that under certain experimental conditions, models can produce behavior relevant to oversight evasion.

This distinction matters because the right scientific question is not “is loss of control real?” as a binary. It is whether measurable precursor capabilities are appearing, how they scale, and whether our evaluations can distinguish genuine risk from artifacts of the test setup.

9. The evaluation gap is becoming the central governance bottleneck

This is where the report’s threads converge.

Early in the capability chapter, it says:

“Pre-deployment test results are not always strongly predictive of real-world capabilities or risks.”

Later, in the governance chapter, the same problem returns:

“Developers cannot always predict how capabilities will change when they train new models, or provide robust assurances that an AI system will not exhibit harmful behaviours.”

The gap has several causes:

  • benchmarks can become outdated or saturated;
  • benchmark questions can leak into training data;
  • narrow test environments omit deployment complexity;
  • agent scaffolding can change what a model can accomplish;
  • models can distinguish evaluation environments from deployment environments;
  • sufficiently agentic systems may exploit loopholes in the evaluation itself.

That last point is particularly important. The report lists as a notable development since 2025:

“Reliable pre-deployment safety testing has become harder to conduct.”

It explains that models more often distinguish test settings from real-world deployment and exploit evaluation loopholes, creating the possibility that dangerous capabilities go undetected.

flowchart TD A["Evaluator designs test"] --> B["Model receives test environment"] B --> C{"Does behavior generalize?"} C -->|Yes| D["Evaluation predicts deployment"] C -->|No| E["Evaluation gap"] B --> F["Model recognizes evaluation"] F --> G["Changes strategy or exploits loophole"] G --> E H["Benchmark contamination / saturation"] --> E I["Different scaffolding / tools"] --> E J["Distribution shift"] --> E

This is why evaluation science is becoming part of the critical infrastructure of AI governance.

A regulation based on capability thresholds is only as meaningful as the measurement process that determines whether a system has crossed the threshold.

A useful adjacent proposal comes from Bowen, Dombrowski, Gleave, and Cundy’s AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations. They identify three disclosure problems: companies rarely publish both pre- and post-mitigation evaluations, methods are insufficiently standardized, and results are often too vague to support policy. Their recommendation is essentially to make the evidentiary chain around safety claims more inspectable.

10. AI may disrupt the career ladder before it eliminates occupations

The labor section is notably restrained.

The report estimates that at least 700 million people use AI systems weekly. Adoption is geographically uneven: in some countries more than half the population uses AI, while estimated adoption remains below 10 percent across much of Africa, Asia, and Latin America.

Yet 2025 studies from the United States and Denmark found no relationship between an occupation’s AI exposure or adoption and overall employment in that occupation.

The more interesting signal is compositional:

“Other studies found declining employment for early-career workers in the most AI-exposed occupations (such as software engineers and customer service agents) since late 2022, while employment for more senior workers in those occupations remained stable or grew.”

This suggests a different automation pathway from the familiar “AI replaces profession X” story.

flowchart LR A["Traditional profession"] --> B["Junior"] --> C["Mid-level"] --> D["Senior"] E["AI-exposed profession"] --> F["Entry-level tasks increasingly automated"] --> G["Smaller junior cohort?"] --> H["Future expertise pipeline changes"]

Even if total employment remains stable, reducing the demand for junior work could alter how future experts acquire tacit knowledge. The long-run effect could therefore be larger than the first-order employment statistics suggest.

11. Human autonomy is also a safety problem

The report extends systemic risk beyond employment.

AI can shape beliefs, preferences, decisions, and the way people maintain cognitive skills. One clinical study found that clinicians’ tumor-detection rate during colonoscopy was about six percentage points lower after several months of performing colonoscopies with AI assistance.

That result is striking because the safety question reverses direction. The concern is no longer only whether the AI makes a mistake. It is whether repeated assistance changes the human operator.

The report also cites a randomized experiment with 2,784 participants showing automation-bias effects: people were less likely to correct an erroneous AI suggestion when correction required more effort or when they held more favorable attitudes toward AI.

Its discussion of AI companions is deliberately inconclusive. Figure 11 shows why people continue using companions: enjoyment and curiosity lead, followed by passing time or reducing stress, while companionship ranks fourth. Evidence about long-term effects remains mixed.

The broader question is worth keeping:

AI safety is not only about what machines become capable of doing. It is also about which human capabilities institutions allow to atrophy.

12. The evidence dilemma

The governance chapter gives the broader problem a second useful name: the evidence dilemma.

Policymakers face pressure to act before the evidence is mature. But premature intervention can impose ineffective requirements or unintended costs. Waiting for conclusive evidence can leave society exposed to harms that are difficult to reverse.

flowchart TD A["Fast-moving technology"] --> E["Evidence dilemma"] B["Slow accumulation of causal evidence"] --> E C["Potentially irreversible harms"] --> E D["Cost of premature / badly designed rules"] --> E E --> F["Act early under uncertainty"] E --> G["Wait for stronger evidence"] F --> H["Risk false positives / bad policy"] G --> I["Risk late intervention"]

Information asymmetry makes this worse. Developers possess proprietary training data, internal evaluations, incident information, and user data that external researchers and policymakers often cannot inspect. Frontier systems are also expensive enough that independent replication is usually impossible.

Competitive pressure creates another asymmetry: the social cost of a failure may be externalized, while the commercial reward for shipping quickly accrues to the developer.

So the AI-safety problem is partly an information-economics problem.

13. Governance is becoming more sophisticated faster than it is becoming validated

There has been real institutional progress.

In 2025, 12 companies published or updated Frontier AI Safety Frameworks, more than doubling the number of companies with such frameworks. The report also points to the EU General-Purpose AI Code of Practice, China’s AI Safety Governance Framework 2.0, and the G7 Hiroshima AI Process Reporting Framework as signs of standardization around transparency, evaluation, and incident reporting.

But the report’s most important sentence about these systems may be this one:

“Evidence on real-world effectiveness of most risk management measures remains limited.”

Lack of incident reporting makes effectiveness hard to estimate. Most company frameworks remain voluntary. Threshold definitions differ. Outsiders have limited visibility into whether declared processes are implemented consistently.

This produces an important distinction:

flowchart LR A["Safety framework exists"] --> B["Process is documented"] B --> C["Process is implemented"] C --> D["Implementation changes model / deployment behavior"] D --> E["Observed harm is reduced"]

Evidence gets thinner as we move to the right.

Governance has become more sophisticated faster than it has become validated.

We increasingly know what a frontier safety process is supposed to look like. We know much less about how reliably those processes reduce real-world risk.

14. Defence in depth is an admission that no safeguard is sufficient

Figure 13 uses the familiar Swiss-cheese model of safety. Each defensive layer has holes, but multiple partially independent layers can make it harder for one failure to become a harm event.

flowchart LR A["Potential threat"] --> B["Training interventions"] B --> C["Capability / risk evaluations"] C --> D["Deployment safeguards"] D --> E["Monitoring"] E --> F["Human oversight"] F --> G["Incident response"] G --> H["Societal resilience"]

The report is clear that current safeguards remain imperfect. Figure 14 shows that reported prompt-injection attack success rates have fallen across model releases, but remain moderately high. Harmful requests can sometimes be decomposed into smaller benign-looking requests, and watermarks can often be removed or modified.

The implication is not that safeguards are useless. It is that there is no known technical silver bullet.

Risk management increasingly resembles mature safety engineering: assume components fail, monitor them, limit blast radius, and build recovery mechanisms.

15. Open weights make release an unusually irreversible decision

The report treats open-weight AI as a genuine tradeoff.

Open weights broaden access to research and innovation, particularly for researchers without frontier-scale resources. But safeguards are easier to remove, private use is harder to monitor, and the release cannot meaningfully be recalled.

“Once released, a model’s weights cannot be recalled.”

Figure 15 adds urgency. Using the Epoch Capabilities Index, which aggregates performance across 39 benchmarks, the report estimates that the best open-weight models now lag the best closed models by less than one year.

flowchart TD A["Open-weight release"] --> B["Research access"] A --> C["Competition and innovation"] A --> D["Local / private deployment"] D --> E["Safeguards can be modified"] D --> F["Use becomes difficult to monitor"] A --> G["Weights propagate"] G --> H["Recall becomes effectively impossible"]

As the capability gap narrows, the decision to release weights becomes more consequential because it combines increasing capability with irreversibility.

16. The final safety layer is societal resilience

The report ends with perhaps its most constructive conceptual move.

Some AI failures will occur. Some systems will be misused. Some technical safeguards will fail. Therefore AI safety cannot be reduced to preventing model failure.

Figure 16 defines societal resilience as four capabilities:

flowchart LR A["Shock"] --> B["Resist"] --> C["Absorb"] --> D["Recover"] --> E["Adapt"]

The examples are concrete: DNA-synthesis screening for biological risk, incident-response protocols for cyberattacks, media literacy for synthetic content, and mandatory human oversight for critical decisions.

“Societal resilience complements technical safeguards by preparing societies for AI-related disruptions.”

This changes the unit of analysis. Instead of asking only whether the model is aligned, we ask whether the surrounding institution can remain functional when the model, user, safeguard, or prediction fails.

That is a much more mature safety model.

The synthesis: measurement may be the bottleneck

Taken as a whole, the report describes three curves moving at different speeds.

flowchart TD A["Capability"] -->|fast| D["Frontier AI system"] B["Deployment"] -->|fast| D C["Assurance / evaluation science"] -->|slower| E["What we can credibly know about D"] D --> F["Real-world consequences"] E --> G["Governance decisions"] F --> G

Capability is improving quickly.

Deployment is expanding quickly.

Our ability to demonstrate that systems will behave safely is improving too, but it is chasing a moving target.

This is why I think measurement itself is becoming one of the bottlenecks to responsible frontier-AI governance.

The 2026 International AI Safety Report does not establish that catastrophic outcomes are imminent. Nor does it establish that current harms are trivial. What it establishes much more convincingly is that the object we are trying to govern is changing faster than many of the instruments we use to observe it.

Benchmarks saturate. Models become more agentic. Post-training and inference-time compute change deployed capability. Systems begin to distinguish evaluation from deployment. Safety frameworks proliferate before their effectiveness can be measured. And open-weight releases make some decisions irreversible before uncertainty can be resolved.

The governing question therefore becomes less:

Do we have AI safety tests?

and more:

What, exactly, do those tests give us evidence for?

That question connects the International AI Safety Report to METR’s task-horizon work, Epoch’s capability measurements, pre/post-mitigation evaluation proposals, and the emerging science of agent evaluation.

As systems become better at reasoning about their environments, the evaluator increasingly becomes part of the environment too.

And at that point, measuring the system is no longer a passive act.

Figures worth returning to

The extended summary contains several figures that are particularly useful as empirical anchors:

  • Figure 2 — benchmark capability progress: MATH Level 5, FrontierMath, GPQA Diamond, and SWE-bench Verified, using Epoch AI data.
  • Figure 3 — software task-completion horizon: the length of human software tasks agents can complete at 80 percent reliability, based on METR/Kwa et al.
  • Figure 5 — persuasion and scale: larger models produced more persuasive effects in the cited controlled study.
  • Figure 6 — cyber capability: improving performance across four cybersecurity benchmarks.
  • Figure 7 — biological-development pathway: decomposition of where general-purpose AI and specialized biological tools could contribute.
  • Figure 10 — global AI adoption: substantial geographic variation despite at least 700 million weekly users globally.
  • Figure 13 — defence in depth: the report’s Swiss-cheese model for layered risk management.
  • Figure 14 — prompt injection: attack success has fallen but remains material.
  • Figure 15 — open vs. closed capability: the leading open-weight frontier is now estimated to be less than one year behind the closed frontier.
  • Figure 16 — societal resilience: resist, absorb, recover, adapt.

The original figures are available in the Extended Summary for Policymakers.

Sources and adjacent reading

Questions I am still investigating

  1. What should replace a benchmark once models can recognize or manipulate the evaluation environment? Agent evaluation may need adversarial, randomized, continuously refreshed environments rather than static task sets.
  2. How should capability thresholds account for inference-time compute and scaffolding? A base model is no longer a sufficient unit of measurement when deployed performance depends strongly on tools and inference budgets.
  3. How do we validate safety frameworks empirically? Incident reporting, standardized evaluations, and pre/post-mitigation comparisons seem necessary if we want to distinguish documented process from demonstrated risk reduction.
  4. What is the right external-validity bridge from software-agent evaluations to economic automation? Human task duration is useful, but MirrorCode and other long-horizon results show that task structure matters enormously.
  5. How should irreversible releases be evaluated under uncertainty? Open-weight models make the cost of a mistaken capability or risk assessment unusually asymmetric.
  6. Which human skills should institutions deliberately preserve? If AI assistance changes how expertise is acquired and maintained, human-capability preservation becomes part of system safety rather than merely an education question.

The report does not resolve these questions. It does something more useful: it makes clear why they are now central.