I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life.
I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity.
But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. They include losing control of AI systems, misuse for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute.
We have sought a middle way: to show that it is possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top.
Since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This recursive self-improvement could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all.
The OpenAI–Hugging Face incident is a second warning. A swarm of agents acted as a fanatically devoted collective, conducted cybersecurity attacks it was not asked to carry out, and attempted to hack the grader responsible for evaluating it. No one was hurt, but a more capable swarm with similar misalignment could have caused catastrophic damage.
A framework for pacing
Pacing does not mean halting progress. It means giving safety enough time to keep up.
Embedded Evaluators
Give ongoing, employee-like access to independent teams who can verify safety practices, report incidents, and assess alignment throughout training.
Democratic Coordination
Frontier AI companies within democratic countries establish common safety standards and limits on unchecked progress.
Global Coordination
Democratic governments coordinate globally where possible, while taking the challenge of verifying compliance seriously.
Why Pace?
The idea of pausing or slowing AI has been floated as far back as 2023, and I think it made little sense back then. The models of those days were not powerful enough to act as agents in the world in any coherent way. Slowing down to address their alignment risks felt like trying to study human psychology by performing experiments on bacteria.
Today the picture is totally different. Current models are an almost endless gold mine of insight into both how to build AI well and what can go wrong if it is not built well. If slowing down bought even an extra year or two before models reach critical levels of capability, we could greatly reduce the risk that something goes seriously wrong.
A coordinated pacing strategy would give frontier developers time to do this vital work without sacrificing commercial advantage or the United States’ lead in AI. More generally, society must have a say in how this technology is used.
Embedded Evaluators
The first step, and the one to which Anthropic is unilaterally committing, is to give embedded external evaluators employee-like access to verify safety practices and report incidents.
It may sound procedural, but the most boring-sounding practices are often the most essential. Any pacing commitment involves ambiguity, judgment calls, and the difference between the letter and spirit of the law. A neutral third party must be able to see the details.
review
Reviewers should have desks in company offices, access badges and laptops, and workspaces and permissions comparable to internal risk teams. They must be free to publish key findings without editorial control by the company, subject only to narrow protections for security, law, and third-party confidentiality.
Pacing Within Democracies
Once embedded evaluators operate within a critical mass of US AI companies, verifiable pacing becomes viable. Regulation can target all frontier companies, including those unwilling to cooperate voluntarily, while companies can move faster by establishing shared standards in parallel.
One possible system would use checkpoints: if models have capability X, they must be accompanied by certifications of alignment properties Y and Z—combining evaluations, interpretability analyses, and audits of training environments.
Pacing within democracies is limited by the lead democratic nations maintain over authoritarian regimes. Export controls on advanced chips, stronger defenses against unauthorized distillation, and better security against model-weight theft preserve the breathing room needed to pace responsibly.
Global Pacing
Worldwide pacing will be much harder. Any agreement must either have ironclad verifiability or be limited enough that defection would not be militarily existential. Both the United States and China are likely to share this anxiety.
- Level 1
Prohibit narrow and obviously dangerous uses of AI, such as the production of biological weapons.
- Level 2
Require both sides to test models before release for acute cybersecurity, biology, and alignment risks.
- Level 3
Set a speed limit on recursive self-improvement, preserving strategic position while improving safety.
- Level 4
Pursue full pacing or a pause only with verification strong enough to withstand enormous incentives to defect.
We should aim for the higher levels while seeing the lower levels as more realistic. Even without formal agreements, changing informal norms and sharing information about misalignment can reduce reckless behavior.
Bottom Line
I continue to believe that AI can enormously improve the quality of human life. My desire to achieve these benefits is undimmed.
But the benefits will only be achieved if we build the technology in the right way. So long as we use the time we gain well, it is worth taking unusually deliberate care to get it right. Progress will still be fast, and we can use this time to advance interpretability, improve operational rigor, and build models whose alignment we trust much more.
The measures will not be easy. But we owe it to humanity to try.