Context
Frontier AI governance is not one regulator, one law, or one safety framework. It is a dependency chain. Model developers control training, weights, post-training, deployment architecture, and much of the evidence about their own systems. Evaluators try to measure dangerous capabilities and safeguard failures. Governments interpret those measurements. Standards and corporate policies translate them into thresholds. Regulators, procurement authorities, export-control agencies, courts, and political leaders determine whether anything must actually change.
The articles in Day 2 of the BlueDot Frontier AI Governance course draws a useful distinction between epistemic power and coercive power. Epistemic power is the ability to produce credible knowledge about what a model can do. Coercive power is the ability to compel a developer, cloud provider, deployer, or downstream actor to change behavior. The UK has invested heavily in the first through the AI Security Institute. China has unusually dense machinery for converting policy concepts into standards and administrative requirements. The United States has enormous material leverage through frontier developers, compute, procurement, national-security institutions, and semiconductor controls, but a fragmented federal regulatory structure. The European Union now supplies the clearest example of a statutory bridge from general-purpose AI risk assessment to enforcement.
The resulting picture is not a hierarchy. It is a control system with several partially connected loops. The strongest institutions are not always the ones with the best information. The best evaluators are not always able to compel action. Corporate risk frameworks can define thresholds without external enforcement. International institutions can establish shared evidence without possessing sovereignty.
The central governance question is therefore narrower than “who regulates AI?” but rather when a frontier model crosses a worrying capability threshold, which institution detects it? who decides what response is required? and who has the power to enforce that this response?
Ecosystem map
The map below is split into vertical pieces that show: 1) who builds capability, 2) who measures it, 3) who governs it nationally, 4) how do those jurisdictions connect internationally, and 5) where the control loop fails.
1. Capability production and corporate controls
2. Measurement and epistemic infrastructure
Evaluators depend on model access, disclosures, APIs, filings, or negotiated testing arrangements with the labs above.
3A. United States
3B. United Kingdom
3C. China
3D. European Union
4. International coordination
5. Critical control loop and failure points
This is the hinge of the whole map: detection only matters if it can compel mitigation.
The governing chain
The map reduces to six functions:
| Function | Main actors | What they can actually do | Main dependency |
|---|---|---|---|
| Create capability | Frontier labs, cloud providers, chip firms, investors | Train, fine-tune, secure, release, withhold, or open-weight models | Compute, capital, talent |
| Measure | AISI, CAISI, METR, labs, researchers, standards bodies | Detect capabilities, vulnerabilities, control failures, misuse pathways | Access to models, methods, test validity |
| Interpret | Corporate safety committees, governments, intelligence agencies, regulators, scientific bodies | Decide whether evidence crosses a threshold | Credible evaluations and risk models |
| Set obligations | Legislatures, regulators, ministries, standards bodies, corporate frameworks | Define what must or should happen when conditions are met | Jurisdiction and institutional authority |
| Enforce | Regulators, courts, export-control authorities, procurement systems, administrative agencies | Fine, restrict, condition access, require disclosure, constrain supply chains | Legal authority and political support |
| Coordinate | NAAIMES, International AI Safety Report, OECD, summits, diplomatic networks | Share methods, evidence, norms, and reporting practices | Voluntary cooperation across sovereign states |
The key asymmetry is visible here. Frontier developers often have both the earliest information and immediate operational control. External evaluators can improve the quality of the evidence, but usually do not possess direct authority over release. Governments may possess authority but lack timely access to the evidence needed to use it well.
The International AI Safety Report 2026 describes the institutional version of this problem as an evidence dilemma. It identifies four recurring obstacles: gaps in scientific understanding, information asymmetries, market failures, and institutional-design and coordination problems. It also describes an evaluation gap because pre-deployment results do not reliably predict real-world performance. My separate note on the report tracks that measurement problem in more detail.
Corporate frontier safety policies are becoming a proto-regulatory layer
METR’s Common Elements of Frontier AI Safety Policies identified twelve companies that had published frontier safety policies by December 2025: Anthropic, OpenAI, Google DeepMind, Magic, Naver, Meta, G42, Cohere, Microsoft, Amazon, xAI, and NVIDIA. These policies commonly link capability evaluations to escalating safeguards, information-security measures, deployment controls, and internal accountability.
This is governance even when it is voluntary. A policy can change internal release decisions, force evaluations before deployment, create documentation, and establish named thresholds. Anthropic’s Responsible Scaling Policy, for example, is now at version 3.4 as of July 8, 2026 and uses Risk Reports, capability thresholds, external review provisions, and a Frontier Safety Roadmap.
The weakness is not that voluntary policies do nothing. The weakness is that the developer designs much of the measurement process, controls much of the underlying information, interprets its own policy, and remains exposed to commercial competition. This is why the corporate layer matters but cannot be the entire governance system.
The triage note treats these frameworks as one of the few mechanisms that can operate at roughly the same speed as model development. The unresolved question is how to make their useful parts legible and enforceable without freezing an immature technical standard into law.
United States: material leverage without one frontier regulator
The current US system is powerful but distributed. The White House AI Action Plan organizes federal policy around accelerating innovation, building AI infrastructure, and international diplomacy and security. This is not a comprehensive frontier-safety statute. It places national competitiveness and infrastructure at the center of federal strategy.
The federal government nevertheless controls several consequential mechanisms. The Commerce Department and Bureau of Industry and Security can constrain access to advanced computing and semiconductor technology. NIST’s Center for AI Standards and Innovation conducts evaluations and measurement work with a national-security orientation. Defense, energy, homeland-security, and intelligence institutions can connect AI capability evidence to existing security authorities. Federal procurement can condition access to major government markets.
CAISI is especially important because it connects government measurement capacity to other parts of the state. In May 2026, for example, it published an evaluation of DeepSeek V4 Pro, comparing the model with US and Chinese capability frontiers. It also works jointly with the UK AISI.
At the state level, California’s Transparency in Frontier Artificial Intelligence Act, SB 53, creates a more direct legal connection to the corporate frontier-safety architecture. It requires covered large developers to maintain and publish frontier AI frameworks and creates incident-reporting and catastrophic-risk transparency duties. This is one path by which voluntary corporate practices become statutory expectations.
My current reading is that US power is strongest at the material layer: compute, chips, cloud, procurement, company concentration, and national security. Its weakness is institutional fragmentation. The question is not whether the United States has leverage. It is how consistently that leverage is connected to a shared risk-detection and decision process.
United Kingdom: unusually strong epistemic capacity
The UK AI Security Institute is the clearest example of a government investing directly in the science of frontier-model evaluation. Its Frontier AI Trends Report synthesizes two years of testing across national-security and public-safety domains. AISI says its mission is to equip governments with empirical understanding of advanced AI, and it has built a technical staff with access to leading models.
AISI’s strength is also its institutional limit. It is not the general-purpose regulator of frontier developers. Evidence has to travel from the evaluator into DSIT, national-security bodies, sector regulators, company processes, or future legislation before it becomes compulsory action.
This is the central tension in The UK’s AI Governance Bet: Britain may have a comparative advantage in producing trusted information about advanced systems, but measurement becomes geopolitical power only when other actors rely on it and when governments have credible pathways for acting on the results.
The UK also coordinates the International Network for Advanced AI Measurement, Evaluation and Science, or NAAIMES. The network includes Australia, Canada, the EU, France, Japan, Kenya, South Korea, Singapore, the UK, and the US. Its 2026 work is focused on evaluation best practice and comparability. This is an important international function, but it remains a measurement network rather than a supranational regulator.
China: standards and administrative machinery
China’s system is structured differently. The Cyberspace Administration of China sits inside a broader party-state apparatus that combines internet governance, data governance, content controls, cybersecurity, industrial strategy, and standards development. The AI Safety Governance Framework 2.0, released in September 2025 under CAC guidance, expands the official risk vocabulary to include open-source model risks, CBRN misuse, labor impacts, risk grading, and loss of human control.
The framework itself is not binding law. The important mechanism is translation. TC260 and CNCERT-CC connect technical policy work to a standards pipeline, while ministries and CAC can turn parts of that pipeline into administrative requirements. China had already used this style of governance for recommendation algorithms, deep synthesis, generative AI services, filings, and security assessments. In 2025, the AI-generated content labeling measures added another direct compliance layer.
Carnegie’s Matt Sheehan and Scott Singer argue in How China Views AI Risks and What to Do About Them that Framework 2.0 is best read as a potential roadmap for standards and later regulation rather than as a stand-alone law. My China governance note develops the same point: China currently has a mature, coercive system for information and platform governance alongside a younger frontier-safety system focused on evaluations, capability risks, open-weight models, and control.
China’s strength is implementation density. Its weakness, relative to the UK and US frontier-evaluation ecosystem, is the maturity and transparency of independent frontier measurement. If China closes that measurement gap while retaining its standards-to-administration pipeline, it could build a particularly tight connection between technical risk classification and state action.
The European Union now matters as the enforcement comparator
The EU is not the geographic focus of the three country notes, but leaving it out would distort the current control map. The AI Act’s general-purpose AI obligations have applied since August 2, 2025. As of August 2, 2026, the European Commission’s enforcement powers for those obligations are in effect, including fines.
The European Commission’s GPAI guidance requires providers to navigate documentation and transparency duties, while providers of GPAI models with systemic risk face additional obligations around risk assessment and mitigation, incident reporting, and cybersecurity. The GPAI Code of Practice is voluntary as a compliance pathway, but the underlying AI Act duties are legal obligations.
This distinction matters. A voluntary code can be nested inside a binding statutory regime. That creates a possible template for frontier governance: technical practices remain adaptable, while the duty to maintain an adequate risk-management process is enforceable.
International institutions: shared epistemology, limited sovereignty
The International AI Safety Report 2026 is backed by more than 30 countries and international organizations and was produced with guidance from more than 100 experts. It does not make policy recommendations. Its governance function is to establish a shared scientific baseline.
NAAIMES performs a related function for measurement methods. The OECD Hiroshima AI Process Reporting Framework 2.0, launched in May 2026, supplies a common voluntary structure for organizations to report risk-management practices. The Bletchley and Seoul summit processes helped normalize frontier-model evaluation and company safety commitments.
These mechanisms matter because cross-border comparability reduces the cost of national action. A regulator does not need to invent an entire evaluation vocabulary from scratch if governments and evaluators already agree on measurement methods, incident categories, or risk-reporting structures.
But none of these institutions has general authority to stop a model release across jurisdictions. International governance is currently strongest at shared epistemology: common evidence, common measurement methods, common reporting structures, and common vocabulary. It is weak at shared sovereignty.
Power map
I use this qualitative matrix as a diagnostic, not a scorecard. “High” means the actor has substantial direct leverage in that category, not that the institution is effective in every case.
| Actor / mechanism | Model information | Evaluation capacity | Rule-setting | Direct enforcement | Cross-border leverage | Main weakness |
|---|---|---|---|---|---|---|
| Frontier developers | High | High | High internally | High internally | High through products | Conflict of interest and competition |
| UK AISI | High when access is granted | High | Low | Low | High epistemic | Depends on others for coercive action |
| US CAISI / NIST | Medium to high | High | Medium through standards | Low directly | Medium to high | Fragmented federal pathway from evaluation to action |
| US BIS / national-security state | Medium | Medium | High | High | Very high | Tools are optimized for security and supply-chain leverage, not general AI risk regulation |
| China CAC / standards system | High through filings and administrative authority | Growing | High | High | Medium | Frontier evaluation ecosystem is less transparent and still maturing |
| EU AI Office / AI Act | High by legal disclosure duties | Growing | High | High | High through market access | New regime, effectiveness still unproven |
| Independent evaluators | Low to medium | High in specialties | Low | None | Medium epistemic | Access and funding dependence |
| NAAIMES | Shared member evidence | High as a network | Soft | None | High epistemic | No supranational enforcement |
| International AI Safety Report | Aggregated evidence | Scientific synthesis | None | None | High agenda-setting | Deliberately non-prescriptive |
| OECD HAIP | Voluntary disclosures | Low | Soft norms | None | Medium | Reporting without enforcement |
The failure points I would watch
1. Information asymmetry
The most capable developers know more than governments about training runs, internal evaluations, deployment incidents, and model-security failures. The International AI Safety Report treats this as a core institutional challenge. Incident reporting, statutory documentation, evaluator access agreements, whistleblower protections, and secure government information-sharing mechanisms all attack the same problem from different directions.
2. The evaluation gap
A test can be methodologically sound and still fail to predict deployment behavior. Models can recognize evaluation environments, exploit benchmark loopholes, perform differently under scaffolding, or acquire capabilities that are difficult to elicit. This is why I treat evaluation institutions as necessary but not sufficient. See AI Capabilities Are Accelerating Faster Than Our Ability to Measure Their Risks and the related METR work on evaluation integrity.
3. The authority gap
This is the most important arrow in the diagram. Suppose AISI, CAISI, a company, or an independent evaluator finds that a model can materially increase biological-weapons capability or autonomously execute dangerous cyber operations. What exactly happens next?
Corporate frameworks may prescribe additional safeguards. California law may trigger disclosure duties. EU law may create statutory risk-management obligations. National-security bodies may have relevant authorities. But there is no universal rule saying that a validated capability threshold automatically triggers a globally recognized deployment restriction.
Measurement -> threshold -> mandatory action remains the weak connection.
4. Competition pressure
Miles Brundage’s triage argument is partly an institutional-capacity argument, but it is also an incentive argument. Governments and firms are making decisions while believing that competitors may move quickly. A safety mechanism that works only when no actor feels strategically pressured is not robust enough for the frontier.
5. Open-weight irreversibility
Once model weights are broadly released, many downstream controls become harder or impossible to impose. China’s Framework 2.0 explicitly recognizes defect propagation and malicious-model risks from open-source systems. The International AI Safety Report also treats open-weight models as a distinct risk-management problem. Governance therefore has to distinguish between reversible deployment decisions and irreversible proliferation decisions.
6. Cross-border mismatch
Models are developed in one jurisdiction, trained on hardware sourced through another, served globally through cloud infrastructure, fine-tuned elsewhere, and evaluated by institutions with different authorities. The governance system is therefore only as strong as its interfaces. Shared evaluation methods help. So do reporting standards. Neither solves jurisdictional arbitrage by itself.
What I think the map says
The field often discusses AI governance as if the missing object were a single law. I do not think that is the right abstraction.
The missing object is a reliable control loop:
Every jurisdiction has pieces of this loop. None has solved all of it.
The UK has invested heavily in observe and measure. China is strong at decide and implement once the party-state establishes a policy direction. The United States has exceptional material leverage over compute, firms, procurement, and security infrastructure, but its decision pathways are distributed across institutions. The EU has built the most explicit legal bridge between systemic-risk duties and enforcement, but its frontier-specific technical capacity and the practical effects of enforcement are still being tested.
International institutions mostly improve observe, measure, and interpret. That is not trivial. Shared science can reduce disagreement about the object being governed. It can also make national action interoperable. But it cannot substitute for institutions that possess legitimate authority to act.
This is why the triage note changes how I read the ecosystem. If governance capacity is scarce and political windows are short, the highest-leverage investments may be the interfaces between layers: evaluator access, mandatory incident reporting, shared threshold definitions, secure government information channels, pre-specified response options, and institutions that can convert a credible warning into action without designing the entire policy response from scratch during a crisis.
That is the part of the map I would keep updating.
References
Related notes
- AI Governance in Triage Mode: What Do We Protect First?
- AI Capabilities Are Accelerating Faster Than Our Ability to Measure Their Risks
- China’s AI Governance Stack: Standards, State Capacity, and the Expanding Definition of Safety
- The UK’s AI Governance Bet: Can Measurement Become Geopolitical Power?
- When Is a Capability Truly Worrying?
- AI Safety Evaluations: CSET Explainer
- How Do We Measure an AI Agent That Can Reason About the Test?
Primary and external sources
- International AI Safety Report, International AI Safety Report 2026, February 3, 2026.
- METR, Common Elements of Frontier AI Safety Policies, December 16, 2025.
- Miles Brundage, We’re in Triage Mode for AI Policy, February 22, 2026.
- UK AI Security Institute, Frontier AI Trends Report, December 2025.
- UK AI Security Institute, International evaluation best practice and open questions in AI measurement, July 23, 2026.
- US NIST, Center for AI Standards and Innovation.
- US NIST, CAISI Evaluation of DeepSeek V4 Pro, May 1, 2026.
- White House, America’s AI Action Plan, July 23, 2025.
- California Legislature, SB 53: Transparency in Frontier Artificial Intelligence Act.
- Cyberspace Administration of China, AI Safety Governance Framework 2.0, September 15, 2025.
- Cyberspace Administration of China, Measures for Labeling AI-Generated and Synthetic Content, March 2025.
- Matt Sheehan and Scott Singer, Carnegie Endowment, How China Views AI Risks and What to Do About Them, October 16, 2025.
- European Commission, Guidelines for providers of general-purpose AI models, updated April 28, 2026.
- European Commission, General-Purpose AI Code of Practice.
- OECD, Hiroshima AI Process Reporting Framework 2.0, May 29, 2026.
- Anthropic, Responsible Scaling Policy, version 3.4 effective July 8, 2026.