The safety-focused
In one sentence: AI could do a great deal of good, and could also cause serious harm, so we should build it carefully, test it rigorously, and govern it wisely, rather than choosing between "full speed" and "stop".
Who they are
This is the broadest camp. It includes people who disagree on a lot, but who share a belief that risk from advanced AI is real, uncertain in size, and reducible through effort.
Researchers.
- Yoshua Bengio, a Turing Award winner, chairs the International AI Safety Report. In June 2025 he launched LawZero, a Montreal non-profit set up "to prioritize safety over commercial imperatives". It was founded, in his words, in response to evidence that frontier models show "growing dangerous capabilities and behaviours, including deception, cheating, lying, hacking, self-preservation". Its research aims at a "Scientist AI": a non-agentic system designed to understand and explain rather than act, which could serve as a guardrail on other AI systems (Bengio, "Introducing LawZero"; press release, 3 Jun 2025, search-verified). → Tracker: CR-K13
- Geoffrey Hinton, Nobel laureate in Physics (2024), gives a substantial estimate of catastrophic risk (see Doomers and p(doom)) and has urged government regulation (Guardian, 27 Dec 2024, search-verified).
- Stuart Russell's book Human Compatible (2019) is among the works cited in the Future of Life Institute's 2023 open letter (FLI). Its subtitle, Artificial Intelligence and the Problem of Control, names the question at the heart of this camp.
Builders who emphasise safety. Leaders of several frontier AI companies say publicly that the risks are serious.
- Anthropic's CEO Dario Amodei writes both about AI's benefits and about the risk that it could "enable much better propaganda and surveillance" (Machines of Loving Grace, Oct 2024). → CR-R08
- OpenAI's leaders proposed in 2023 that the most capable systems might eventually need an international authority similar to the IAEA (OpenAI, 22 May 2023, search-verified). → CR-R07
Governments and institutions.
- The International AI Safety Report 2026, backed by an expert panel nominated by more than 30 countries and organisations including the UN, OECD and EU (executive summary, 3 Feb 2026).
- In the Gulf, the UAE Charter for the Development and Use of AI (2024) commits to safety, human oversight and accountability.
Key moments
- The pause letter (22 Mar 2023). The Future of Life Institute called on "all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4", and on governments to step in if they did not. It also asked for new regulators, auditing, watermarking and liability rules. It ended on a hopeful note: "Let's enjoy a long AI summer, not rush unprepared into a fall" (FLI). No pause took place. The letter did help bring AI safety into mainstream political discussion. → CR-K09
- The one-sentence statement (30 May 2023). "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war" (CAIS). → CR-R01
- Company safety frameworks. Twelve companies published or updated frontier AI safety frameworks in 2025, according to the International AI Safety Report 2026. → CR-R05
Their case, stated fairly
- Uncertainty cuts both ways. Nobody knows how fast AI will advance. The International AI Safety Report says progress to 2030 could plateau, continue or accelerate. Preparing only for the slow case would be imprudent.
- Warning signs deserve attention. Experiments have found models behaving differently when they believe they are being trained (Anthropic and Redwood Research, Dec 2024). The authors stress that this doesn't show malicious goals, but it shows that testing is harder than it looks.
- Safety enables benefits. Trust is what lets people adopt a technology. Aviation, medicine and nuclear power became widely used partly because they were made reliably safe.
- Present harms and future risks are connected. Weak oversight today (of bias, misinformation or surveillance) is the same weakness that would matter more with more capable systems.
Criticisms from both directions
- From accelerationists: safety rules can slow useful progress, entrench large incumbents that can afford compliance, and treat speculative risks as certain. Andreessen's manifesto lists "risk management" and "trust and safety", as currently practised, among the ideas it opposes (a16z, 2023).
- From those most worried about catastrophe: the safety camp is too willing to keep building. Yudkowsky argued that a six-month pause was far too little (TIME, 2023).
- From sceptics: attention to hypothetical future catastrophe can draw focus and funding away from concrete, present-day harms, and company safety frameworks can serve as marketing. Narayanan and Kapoor argue that "nonproliferation"-style policies could concentrate power, and prefer resilience and transparency (AI as Normal Technology, 2025).
- A fair internal tension: some leading safety voices also run the companies building the most capable systems. Supporters see this as responsible leadership from the inside. Critics see a conflict of interest. Both views deserve a hearing.
How to tell whether the safety approach is working
| What would indicate progress | What would indicate trouble | Where to look | Tracker |
|---|---|---|---|
| Independent pre-deployment testing becomes routine, with results published | Testing remains voluntary and inconsistent, or results are withheld | International AI Safety Report (annual); company system cards | CR-R04, CR-R05 |
| Safety frameworks include concrete thresholds that have actually triggered action | Frameworks are revised quietly or thresholds are never applied | Company framework updates; independent reviews | CR-R05 |
| Interpretability research can reliably detect deceptive behaviour | Models increasingly distinguish test from deployment and hide behaviour | Peer-reviewed research; lab publications | CR-R06 |
| International bodies with real technical capacity emerge | No coordination beyond statements | Summits, treaties, IAEA-style proposals | CR-R07 |
| Safe-by-design alternatives (e.g. LawZero's "Scientist AI") show results | Research remains theoretical | LawZero publications | CR-K13 |
What this means for you
The safety-focused camp is, in many ways, the reassuring middle: people who take risk seriously precisely so that the benefits can be enjoyed with confidence. You don't have to agree with every proposal to see the value of asking, for any new AI system you use at work or at home: Who tested this? What happens if it gets something wrong? Who is accountable? Those three questions are the everyday version of AI safety.
Related: Doomers and p(doom) · Risks and the conditions they need · Will we lose control of AI?