The safety-focused

In one sentence: AI could do a great deal of good, and could also cause serious harm, so we should build it carefully, test it rigorously, and govern it wisely, rather than choosing between "full speed" and "stop".

Who they are

This is the broadest camp. It includes people who disagree on a lot, but who share a belief that risk from advanced AI is real, uncertain in size, and reducible through effort.

Researchers.

  • Yoshua Bengio, a Turing Award winner, chairs the International AI Safety Report. In June 2025 he launched LawZero, a Montreal non-profit set up "to prioritize safety over commercial imperatives". It was founded, in his words, in response to evidence that frontier models show "growing dangerous capabilities and behaviours, including deception, cheating, lying, hacking, self-preservation". Its research aims at a "Scientist AI": a non-agentic system designed to understand and explain rather than act, which could serve as a guardrail on other AI systems (Bengio, "Introducing LawZero"; press release, 3 Jun 2025, search-verified). → Tracker: CR-K13
  • Geoffrey Hinton, Nobel laureate in Physics (2024), gives a substantial estimate of catastrophic risk (see Doomers and p(doom)) and has urged government regulation (Guardian, 27 Dec 2024, search-verified).
  • Stuart Russell's book Human Compatible (2019) is among the works cited in the Future of Life Institute's 2023 open letter (FLI). Its subtitle, Artificial Intelligence and the Problem of Control, names the question at the heart of this camp.

Builders who emphasise safety. Leaders of several frontier AI companies say publicly that the risks are serious.

  • Anthropic's CEO Dario Amodei writes both about AI's benefits and about the risk that it could "enable much better propaganda and surveillance" (Machines of Loving Grace, Oct 2024). → CR-R08
  • OpenAI's leaders proposed in 2023 that the most capable systems might eventually need an international authority similar to the IAEA (OpenAI, 22 May 2023, search-verified). → CR-R07

Governments and institutions.

  • The International AI Safety Report 2026, backed by an expert panel nominated by more than 30 countries and organisations including the UN, OECD and EU (executive summary, 3 Feb 2026).
  • In the Gulf, the UAE Charter for the Development and Use of AI (2024) commits to safety, human oversight and accountability.

Key moments

  • The pause letter (22 Mar 2023). The Future of Life Institute called on "all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4", and on governments to step in if they did not. It also asked for new regulators, auditing, watermarking and liability rules. It ended on a hopeful note: "Let's enjoy a long AI summer, not rush unprepared into a fall" (FLI). No pause took place. The letter did help bring AI safety into mainstream political discussion. → CR-K09
  • The one-sentence statement (30 May 2023). "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war" (CAIS). → CR-R01
  • Company safety frameworks. Twelve companies published or updated frontier AI safety frameworks in 2025, according to the International AI Safety Report 2026. → CR-R05

Their case, stated fairly

  1. Uncertainty cuts both ways. Nobody knows how fast AI will advance. The International AI Safety Report says progress to 2030 could plateau, continue or accelerate. Preparing only for the slow case would be imprudent.
  2. Warning signs deserve attention. Experiments have found models behaving differently when they believe they are being trained (Anthropic and Redwood Research, Dec 2024). The authors stress that this doesn't show malicious goals, but it shows that testing is harder than it looks.
  3. Safety enables benefits. Trust is what lets people adopt a technology. Aviation, medicine and nuclear power became widely used partly because they were made reliably safe.
  4. Present harms and future risks are connected. Weak oversight today (of bias, misinformation or surveillance) is the same weakness that would matter more with more capable systems.

Criticisms from both directions

  • From accelerationists: safety rules can slow useful progress, entrench large incumbents that can afford compliance, and treat speculative risks as certain. Andreessen's manifesto lists "risk management" and "trust and safety", as currently practised, among the ideas it opposes (a16z, 2023).
  • From those most worried about catastrophe: the safety camp is too willing to keep building. Yudkowsky argued that a six-month pause was far too little (TIME, 2023).
  • From sceptics: attention to hypothetical future catastrophe can draw focus and funding away from concrete, present-day harms, and company safety frameworks can serve as marketing. Narayanan and Kapoor argue that "nonproliferation"-style policies could concentrate power, and prefer resilience and transparency (AI as Normal Technology, 2025).
  • A fair internal tension: some leading safety voices also run the companies building the most capable systems. Supporters see this as responsible leadership from the inside. Critics see a conflict of interest. Both views deserve a hearing.

How to tell whether the safety approach is working

What would indicate progressWhat would indicate troubleWhere to lookTracker
Independent pre-deployment testing becomes routine, with results publishedTesting remains voluntary and inconsistent, or results are withheldInternational AI Safety Report (annual); company system cardsCR-R04, CR-R05
Safety frameworks include concrete thresholds that have actually triggered actionFrameworks are revised quietly or thresholds are never appliedCompany framework updates; independent reviewsCR-R05
Interpretability research can reliably detect deceptive behaviourModels increasingly distinguish test from deployment and hide behaviourPeer-reviewed research; lab publicationsCR-R06
International bodies with real technical capacity emergeNo coordination beyond statementsSummits, treaties, IAEA-style proposalsCR-R07
Safe-by-design alternatives (e.g. LawZero's "Scientist AI") show resultsResearch remains theoreticalLawZero publicationsCR-K13

What this means for you

The safety-focused camp is, in many ways, the reassuring middle: people who take risk seriously precisely so that the benefits can be enjoyed with confidence. You don't have to agree with every proposal to see the value of asking, for any new AI system you use at work or at home: Who tested this? What happens if it gets something wrong? Who is accountable? Those three questions are the everyday version of AI safety.

Related: Doomers and p(doom) · Risks and the conditions they need · Will we lose control of AI?