KnowledgeInSight
AI Literacy
0% of Course 3 complete

Module 4 · Lesson 2

Governance: Company self-governance and its critics

You'll look at the safety commitments AI companies have made, how anyone checks them, and two opposite criticisms: that they're too weak, and that safety talk protects large firms from competition. You'll be able to judge a company commitment by what it promises and who verifies it.

What you will be able to do

  • Evaluate what company safety commitments promise, how they are checked, and where critics say they fall short.

0% of this lesson · 9 items · 1h 3m total · 48m without the optional journal

Contents of this lesson9 items
  1. ReadingPromises Without a Regulator: What Voluntary Safety Commitments Are3 min
  2. ReadingFrontier Safety Frameworks: Capability Thresholds and If-Then Commitments4 min
  3. ReadingHow Commitments Are Checked: Government Institutes, Outside Evaluators, and Whistleblowers4 min
  4. ReadingThe Criticism That Commitments Are Too Weak or Are Weakening4 min
  5. ReadingThe Criticism That Safety Talk Serves Incumbents: Regulatory Capture and the Open-Model Argument4 min
  6. Guided ReadingGuided Close Reading: The Frontier AI Safety Commitments Signed in Seoul7 min
  7. Guided ConversationHear Both Critiques of Self-Governance12 min
  8. Journal · optionalWhat Would Make a Promise Credible?15 min
  9. Knowledge CheckCompany self-governance and its critics10 min

Reading 3 min

Promises Without a Regulator: What Voluntary Safety Commitments Are

This content reflects the field as of October 2026.

When a new medicine or a new aircraft reaches the public, a government agency has usually examined it first. News coverage of AI often assumes something similar happens before a powerful new model is released. In most countries it doesn't.

As of October 2026, the main limits on how the most capable AI systems are built and released are limits the companies wrote for themselves. A few laws now touch these systems, and most of those require companies to publish or report things. The decisions that matter most for safety, such as how much testing is enough, what result would stop a release, and what protections a model needs, are mostly set by each company's own policy.

Those policies aren't purely private. In May 2024, at a summit in Seoul hosted by the governments of the United Kingdom and South Korea, sixteen organizations agreed to a common list called the Frontier AI Safety Commitments. Four more joined later. The signers include Amazon, Anthropic, Google, Meta, Microsoft, OpenAI, and xAI from the United States, along with companies based in Canada, China, France, South Korea, and the United Arab Emirates (Department for Science, Innovation and Technology 2024).

The document describes itself with care. The organizations "undertake to develop and deploy their frontier AI models and systems responsibly," following what the text calls "voluntary commitments" (Department for Science, Innovation and Technology 2024). They promise to assess the risks of their models, to set thresholds at which a risk would be intolerable, and to publish a framework explaining how they do both.

The word "voluntary" carries the weight. Two governments announced the list, and neither enforces it. The text names no inspector and no penalty.

People draw opposite conclusions from this arrangement.

  • Some see promises with nothing behind them. On this view a company will keep a commitment until it becomes expensive, and then rewrite it.
  • Some see a sensible first step. The technology is changing quickly, governments lack the expertise to write detailed rules, and public commitments at least give outsiders something to hold companies to.
  • Some see a different problem. They argue that the largest companies talk about danger because rules built around their own safety practices would be hard for smaller rivals to meet.

Each of these views is held by serious people, and each group has interests of its own. Companies gain from being trusted without being regulated. Advocacy groups that warn about AI gain attention and funding when the warnings are believed. Firms that sell openly released models gain when restrictions are seen as self-serving.

You don't need to settle that dispute to read a commitment well. Three questions work on any promise a company makes about AI safety, whether it appears in a summit document, a company policy, or a press release.

  1. What exactly is promised? Look for the verb and for the conditions attached.
  2. Who checks? Someone outside the company either can see whether the promise was kept, or can't.
  3. What happens if it's broken? There may be a penalty, a legal duty, or nothing.

References

  • Department for Science, Innovation and Technology. 2024. "Frontier AI Safety Commitments, AI Seoul Summit 2024." May 21, 2024.

Report an issue with this item

Reading 4 min

Frontier Safety Frameworks: Capability Thresholds and If-Then Commitments

This content reflects the field as of October 2026.

Introduction

AI companies say they test their most powerful models for dangerous abilities before release. A reasonable follow-up is what they've promised to do if a test comes back with a worrying result.

This reading explains the documents in which several developers answer that question, how three of them are built, and what an independent comparison found across twelve.

What the Seoul Commitments Asked For

In May 2024 sixteen organizations, later twenty, agreed to the Frontier AI Safety Commitments at a summit in Seoul. They promised to assess the risks of their most capable models, to set out thresholds at which severe risks "would be deemed intolerable," to be transparent about how they do this, and to publish "a safety framework focused on severe risks" (Department for Science, Innovation and Technology 2024).

A frontier safety framework is a document in which an AI developer states how it will test its most capable models for severe risks and what it will do about the results. Different companies use different titles for it.

How a Framework Works

Most frameworks share one mechanism. A capability threshold is a specified level of ability at which a model is judged to pose a serious enough risk to require stronger protections. A safeguard is a protection a developer puts in place, such as security against theft of a model or limits on what the model will do for users.

The two are joined in an if-then commitment: a promise that if a model reaches a stated threshold, then the developer will apply stated safeguards before going further. For example, if a model could give real help to someone trying to make a biological weapon, then it won't be released until protections against that use are in place.

Three companies' frameworks show the pattern.

CompanyFrameworkHow thresholds are definedVersion read here
AnthropicResponsible Scaling PolicyCapability thresholds in areas that include chemical, biological, radiological, and nuclear weapons and the automation of AI research, each tied to required standards of security and deployment safeguardsVersion 3.4, July 2026
Google DeepMindFrontier Safety Framework"Critical Capability Levels," described as levels at which, without mitigation, a model "may pose heightened risk of severe harm"Version 3.0, September 2025, with a later update
OpenAIPreparedness FrameworkTwo levels, "High" and "Critical," in tracked categories: biological and chemical, cybersecurity, and AI self-improvementVersion 2, April 2025

(Anthropic 2026; Google DeepMind 2025; OpenAI 2025c.)

The three differ in vocabulary and in which risks they track. They share a structure: named abilities, tests for them, and protections required once a model shows the ability.

What a Comparison Across Twelve Companies Found

METR, a nonprofit that evaluates AI models, compared the published policies of twelve companies in December 2025. Besides the three above, they included Amazon, Cohere, G42, Magic, Meta, Microsoft, Naver, Nvidia, and xAI (METR 2025).

It found nine elements that recur, among them capability thresholds, security for model weights, conditions for halting deployment and development, and a process for updating the policy. It also found differences. Some policies concentrate on catastrophic risks and others on a wider range of harms, and some lean more heavily on numerical test scores (METR 2025).

METR has evaluated models for some of the companies it compared, which gives it knowledge of their policies and a working relationship with their authors.

What a Framework Is and Isn't

A framework is a company's statement of what it will do. The company writes it, decides whether a threshold has been reached, chooses the safeguards, and can revise the document. Each of the three above has been revised at least once.

One law has changed the picture slightly. California's 2025 statute requires large developers of the most capable models to publish a framework on their websites (Office of the Governor of California 2025). The law makes publication a legal duty in that state. The thresholds and safeguards inside the framework remain the company's own choices.

Some developers also publish a second kind of document, under titles such as "model spec" or "constitution," describing how their models are meant to behave in ordinary use. Those documents concern everyday behavior. Safety frameworks concern rare and severe risks.

Conclusion

As of October 2026, several leading developers publish frameworks that tie stronger safeguards to specified levels of capability. The frameworks share a structure and differ in detail, and an independent comparison has documented both. They are public commitments written and interpreted by the companies that make them, with publication required by law in at least one US state.

Key Terms

  • Frontier safety framework: A document in which an AI developer states how it will test its most capable models for severe risks and what it will do about the results.
  • Capability threshold: A specified level of ability at which a model is judged to pose a serious enough risk to require stronger protections.
  • Safeguard: A protection a developer puts in place, such as security against theft of a model or limits on what the model will do for users.
  • If-then commitment: A promise that if a model reaches a stated threshold, then the developer will apply stated safeguards before going further.

References

  • Anthropic. 2026. "Responsible Scaling Policy." Version 3.4, effective July 8, 2026.
  • Department for Science, Innovation and Technology. 2024. "Frontier AI Safety Commitments, AI Seoul Summit 2024." May 21, 2024.
  • Google DeepMind. 2025. "Frontier Safety Framework." Version 3.0, September 22, 2025.
  • METR. 2025. "Common Elements of Frontier AI Safety Policies." December 2025.
  • Office of the Governor of California. 2025. "Governor Newsom Signs SB 53, Advancing California's World-Leading Artificial Intelligence Industry." September 29, 2025.
  • OpenAI. 2025c. "Our Updated Preparedness Framework." Version 2, April 15, 2025.

Report an issue with this item

Reading 4 min

How Commitments Are Checked: Government Institutes, Outside Evaluators, and Whistleblowers

This content reflects the field as of October 2026.

Introduction

A company can publish a detailed safety policy and follow none of it, and from outside you couldn't tell. Whether a commitment means anything depends on whether someone other than the company can see what the company does.

This reading covers three ways commitments get checked, a 2026 proposal to strengthen one of them, and the weakness they share.

Government Institutes

Verification is checking, by someone other than the party that made a promise, that the promise is being kept. Several governments have set up a safety institute: a government body that tests advanced AI models and studies their risks.

The United Kingdom's is the AI Security Institute, part of the Department for Science, Innovation and Technology. It was founded as the AI Safety Institute and later renamed. It describes its work as "testing leading AI systems before they are released publicly," and says it has "pre-deployment access to leading AI models" through collaboration with AI companies (UK AI Security Institute n.d.).

The US counterpart has changed more. In June 2025 the Commerce Department turned the US AI Safety Institute into the Center for AI Standards and Innovation. The commerce secretary's statement said the center would set up "voluntary agreements" with developers and lead evaluations of "demonstrable risks, such as cybersecurity, biosecurity, and chemical weapons." It also said that "censorship and regulations have been used under the guise of national security" (US Department of Commerce 2025). As of October 3, 2026 the center's official page carries another new name, the Center for Advancing Innovation and Standards for Super Intelligence (National Institute of Standards and Technology n.d.).

Neither institute is a regulator. Neither can block a release.

Outside Evaluators

Third-party evaluation is the testing of a model by an organization that didn't build it. Nonprofits and specialist firms do this work, sometimes at a developer's invitation before a release.

What they can learn depends on access: what an outside party is permitted to see and test, and for how long. An evaluator given a few days with a finished model learns less than one who can see training methods, internal test results, and incident reports.

Employees Who Speak

A whistleblower is an employee or former employee who reports wrongdoing or danger inside an organization to people outside it. Employees see what no outside evaluator sees.

In June 2024 a group of current and former employees of OpenAI and Google DeepMind published an open letter titled "A Right to Warn About Advanced Artificial Intelligence." It asked AI companies to accept four principles ("A Right to Warn" 2024):

  • not to make or enforce agreements that prohibit criticism of the company over risk-related concerns
  • to provide an anonymous process for raising such concerns with the board, regulators, and an independent organization
  • to support a culture of open criticism, with trade secrets protected
  • not to retaliate against employees who share risk-related information publicly after other processes have failed

The letter is a request. Its authors argued that ordinary whistleblower laws cover illegal activity, while many of the risks they worried about weren't yet regulated at all.

A 2026 Proposal

In September 2026 Dario Amodei, chief executive of Anthropic, published an essay proposing "embedded evaluators": outside reviewers placed inside a company with "ongoing, employee-like access," able to check whether the company follows the practices it claims and to publish findings without the company editing them. The essay said Anthropic would do this on its own and urged other companies to follow (Amodei 2026b).

The news site Axios reported that Sam Altman, chief executive of OpenAI, said the same day that his company would give outside evaluators similar access (Berkowitz 2026). Both men lead companies that would gain from being seen as trustworthy, and the essay is its author's account of his own company's intentions. As of October 2026 these are announcements. Whether the access is delivered, and to whom, is something to watch for.

What's Still Missing

The three routes share a weakness. In each, the company controls the door.

RouteWhat it depends on
Government instituteAgreements the company chooses to sign
Outside evaluatorAccess the company chooses to grant
EmployeeContract terms and internal channels the company sets

Access that is granted can be narrowed or withdrawn, and an evaluator who depends on a company's goodwill has a reason to stay on good terms with it. The embedded-evaluator proposal tries to answer this with standing access and publication rights. It would still rest on the company's consent unless a law required it.

Conclusion

As of October 2026, company safety commitments are checked by government institutes with testing agreements, by outside evaluators, and by employees willing to speak. Each route has produced real information. Each depends on access that the companies grant, and proposals to make that access standing or mandatory haven't yet been tested.

Key Terms

  • Verification: Checking, by someone other than the party that made a promise, that the promise is being kept.
  • Safety institute: A government body that tests advanced AI models and studies their risks.
  • Third-party evaluation: The testing of a model by an organization that didn't build it.
  • Access: What an outside party is permitted to see and test, and for how long.
  • Whistleblower: An employee or former employee who reports wrongdoing or danger inside an organization to people outside it.

References

  • Amodei, Dario. 2026b. "We Must Pace the Frontier." September 2026.
  • Berkowitz, Ben. 2026. "Anthropic, OpenAI CEOs Call for Slowdown in AI Development." Axios, September 12, 2026.
  • National Institute of Standards and Technology. n.d. "Center for Advancing Innovation and Standards for Super Intelligence (CAISSI)." Accessed October 3, 2026.
  • "A Right to Warn About Advanced Artificial Intelligence." 2024. Open letter, June 4, 2024.
  • UK AI Security Institute. n.d. "About." Accessed October 3, 2026.
  • US Department of Commerce. 2025. "Statement from U.S. Secretary of Commerce Howard Lutnick on Transforming the U.S. AI Safety Institute into the Pro-Innovation, Pro-Science U.S. Center for AI Standards and Innovation." June 3, 2025.

Report an issue with this item

Reading 4 min

The Criticism That Commitments Are Too Weak or Are Weakening

This content reflects the field as of October 2026.

Introduction

AI companies point to their safety policies as evidence that they can be trusted with powerful systems. One group of critics reads the same policies and sees promises that are thin to begin with and that loosen when keeping them gets costly.

This reading sets out that criticism, the evidence its advocates cite, and the companies' reply. The dispute is unresolved.

The Grades

A voluntary commitment is a promise that an organization makes by its own choice and that no law requires it to keep. The critics' first claim is that the existing ones fall short of the risks the companies themselves describe.

The most systematic version is a safety index: a published scorecard that grades companies on their safety practices using stated criteria. The Future of Life Institute, a nonprofit, issues one. Its summer 2026 edition had a panel of seven researchers grade nine companies in six areas (Future of Life Institute 2026).

CompanyOverall grade
AnthropicC+
Google DeepMindC
OpenAIC
MetaD+
Alibaba CloudD-
Z.aiD-
DeepSeekF
MistralF
xAIF

No company scored above C+. In the area the index calls existential safety, meaning preparation for risks that could threaten humanity as a whole, the report says: "No company exceeds C-; most score D or below." One of its headline findings is that "safety rhetoric outpaces revealed behavior" (Future of Life Institute 2026).

The Future of Life Institute campaigns for stronger regulation of advanced AI, and low grades support its case. The grades are its panel's judgments against its own criteria.

The Changes

The second claim is that commitments have loosened. The critics cite two documented cases.

In February 2026 Anthropic replaced its policy with a rewritten version. TIME magazine reported that the company had dropped the central pledge of its 2023 policy: not to train more capable systems unless it could guarantee in advance that its safety measures were adequate. Under the new version, the magazine reported, the company commits to delay development only if it considers itself the leader of the field and judges the risk of catastrophe to be significant (Perrigo 2026).

OpenAI's framework, as updated in April 2025, contains a clause about rivals. It says: "If another frontier AI developer releases a high-risk system without comparable safeguards, we may adjust our requirements." The clause sets conditions. The company would first confirm that the risks had actually changed, acknowledge the adjustment publicly, and keep safeguards "at a level more protective" (OpenAI 2025c).

The Argument

Critics connect these cases with one idea. Competitive pressure is the push a company feels to match its rivals or lose customers and investment to them. A company that holds back alone bears the cost of caution.

On this argument a voluntary commitment is weakest when a rival is about to get ahead. Chris Painter, policy director at the evaluation nonprofit METR, told TIME that without firm thresholds he worried about a "frog-boiling effect," in which danger rises gradually and no single step sets off an alarm (Perrigo 2026).

The critics' remedy is a binding rule: a requirement set by law, with a penalty for breaking it, that applies to every company alike. If every company must meet the same standard, none loses ground by meeting it. Bill Gates, co-founder of Microsoft, made a version of this argument in a September 2026 podcast interview, saying the industry can't be relied on to regulate itself (Klein 2026c).

The Companies' Reply

The companies give three answers.

  • Frameworks have to change. The science of testing models is young. Jared Kaplan, Anthropic's chief science officer, told TIME: "We felt that it wouldn't actually help anyone for us to stop training AI models." The company's argument was that stopping alone would leave the pace to developers with weaker protections (Perrigo 2026).
  • Changes are made in the open. The revised policies are published with their revision histories, and OpenAI's clause promises public acknowledgment of any adjustment. A later version of Anthropic's policy states that the company remains free to pause development "even if not required" by the policy (Anthropic 2026).
  • Some voluntary restraint has happened. In September 2026, a monthly review of US technology policy reported, OpenAI paused training of its most capable models and said it wouldn't release a new model, citing safety concerns (Lau and others 2026).

Each answer comes from a party with a stake. The companies gain if voluntary commitments are judged adequate.

Conclusion

As of October 2026 it's documented that company safety policies have been revised, and that at least two revisions made a commitment more conditional. It's also documented that policies are published and that one company paused on its own. Whether voluntary commitments can hold under competition is a forecast on both sides. The critics' forecast leads to binding rules, and the companies' leads to continued self-governance with more transparency.

Key Terms

  • Voluntary commitment: A promise that an organization makes by its own choice and that no law requires it to keep.
  • Safety index: A published scorecard that grades companies on their safety practices using stated criteria.
  • Competitive pressure: The push a company feels to match its rivals or lose customers and investment to them.
  • Binding rule: A requirement set by law, with a penalty for breaking it, that applies to every company alike.

References

  • Anthropic. 2026. "Responsible Scaling Policy." Version 3.4, effective July 8, 2026.
  • Future of Life Institute. 2026. AI Safety Index: Summer 2026. July 2026.
  • Klein, Ezra, host. 2026c. "Bill Gates's Blunt Warning on A.I." The Ezra Klein Show, podcast, New York Times, September 29, 2026.
  • Lau, Rachel, Shirley Frame, Justin Hendrix, and Ashley Faler. 2026. "September 2026 US Tech Policy Roundup." Tech Policy Press, October 1, 2026.
  • OpenAI. 2025c. "Our Updated Preparedness Framework." Version 2, April 15, 2025.
  • Perrigo, Billy. 2026. "Exclusive: Anthropic Drops Flagship Safety Pledge." TIME, February 24, 2026.

Report an issue with this item

Reading 4 min

The Criticism That Safety Talk Serves Incumbents: Regulatory Capture and the Open-Model Argument

This content reflects the field as of October 2026.

Introduction

When the head of an AI company warns that AI is dangerous, some listeners hear a responsible warning. Others ask why a company would advertise the dangers of its own product, and suspect that the warning is good for business.

This reading sets out the criticism that safety warnings help large companies hold off competitors, the related argument for openly released models, and the replies. The dispute is unresolved.

Regulatory Capture

Regulatory capture is a situation in which the companies being regulated shape the rules to serve themselves. The rules then protect those companies from competition more than they protect the public.

An incumbent is a company that already holds a strong position in a market. The capture argument says incumbents can afford safety teams, lawyers, and compliance staff. Rules that demand those things cost them little and cost a newcomer a great deal.

Andrew Ng, founder of the education company DeepLearning.AI and an investor in AI startups, has made this argument repeatedly. In March 2026 he wrote of occasions "when big AI companies argue that AI is dangerous to block the free distribution of open source projects that compete with their offerings" (Ng 2026a). In July 2026 he went further, writing that "a meaningful fraction of work on AI safety is no longer about safety but rather aimed at stoking fears to pursue regulatory capture" (Ng 2026c).

Jensen Huang, chief executive of the chip maker Nvidia, argued in a September 2026 podcast interview that warnings of catastrophe frighten the public without cause and that the industry needs no new regulation (Klein 2026b).

The Open-Model Argument

Much of this argument concerns one kind of product. An open-weight model is a model whose parameters have been published for anyone to download. Parameters, also called weights, are the numbers that hold what a model learned in training. Anyone with suitable hardware can run such a model, study it, and modify it.

Advocates make three claims for open release.

  • It spreads capability. Researchers, small companies, and poorer countries can build on a model they couldn't afford to train.
  • It allows inspection. Outsiders can examine an open model directly. Ng wrote that open release "casts sunlight on technology and ultimately makes it safer" (Ng 2026c).
  • It prevents gatekeeping. Gatekeeping is control by a few parties over who may use something and on what terms. Yann LeCun, then chief AI scientist at Meta, said in a 2024 interview: "I see the danger of this concentration of power through proprietary AI systems as a much bigger danger than everything else" (LeCun 2024).

The Reply

The 2026 International AI Safety Report, an expert assessment chaired by the computer scientist Yoshua Bengio, accepts that open-weight models help research and give access to people with fewer resources. It then identifies what is different about them. "Once released, a model's weights cannot be recalled." Their "safeguards are easier to remove," and monitoring is harder because anyone can run them outside controlled settings (Bengio and others 2026, sec. 3.4).

The reply to the capture charge is that a motive doesn't settle whether a claim is true. A company can profit from a warning that is also correct.

One Proposal, Two Readings

A September 2026 episode shows both readings at once. Dario Amodei, chief executive of Anthropic, published an essay arguing that companies and governments should slow the rate at which AI capabilities improve, with outside verification (Amodei 2026b). Sam Altman, chief executive of OpenAI, said he agreed, as did Elon Musk of xAI (Berkowitz 2026).

Supporters read this as leading companies accepting limits on themselves. The technology news site The Register wrote that the proposal reads "like an attempt at 'regulatory capture'" and would give the largest companies a truce in which they needn't compete so fiercely (The Register 2026). Chamath Palihapitiya, a venture capitalist, called the essay an effort to stop open source and concentrate power with its author's company (Berkowitz 2026). Amodei's essay acknowledges that his company's advocacy has led to accusations of "regulatory capture" (Amodei 2026b).

Who Benefits From Each Position

PositionWho holds itWhat they stand to gain
Safety rules are neededCompanies whose leading models are closed, such as Anthropic and OpenAI; safety researchersRules they can meet more easily than smaller rivals; funding and standing for safety work
Safety talk is captureDevelopers and users of open models, such as Meta; chip makers such as Nvidia; investors in startupsFreedom to release and build on open models; more buyers for chips

Every party here argues from interest. That doesn't show that any of them is wrong.

Conclusion

The capture critique holds that warnings about AI danger help incumbents restrict competitors, above all the makers of open-weight models. The reply holds that open release carries risks that can't be reversed, and that a self-interested warning can still be true. As of October 2026 both sides can point to named evidence: experts document the risks of open release, and critics document who would gain from restricting it.

Key Terms

  • Regulatory capture: A situation in which the companies being regulated shape the rules to serve themselves.
  • Incumbent: A company that already holds a strong position in a market.
  • Open-weight model: A model whose parameters have been published for anyone to download.
  • Gatekeeping: Control by a few parties over who may use something and on what terms.

References

  • Amodei, Dario. 2026b. "We Must Pace the Frontier." September 2026.
  • Bengio, Yoshua, and others. 2026. International AI Safety Report 2026: Extended Summary for Policymakers. Published February 3, 2026.
  • Berkowitz, Ben. 2026. "Anthropic, OpenAI CEOs Call for Slowdown in AI Development." Axios, September 12, 2026.
  • Klein, Ezra, host. 2026b. "Jensen Huang Thinks A.I. Alarmism Has Gone Too Far." The Ezra Klein Show, podcast, New York Times, September 23, 2026.
  • LeCun, Yann. 2024. Interview by Lex Fridman. Lex Fridman Podcast no. 416, transcript, March 8, 2024.
  • Ng, Andrew. 2026a. "How Anti-AI Propaganda Hurts the Public." The Batch, DeepLearning.AI, March 27, 2026.
  • Ng, Andrew. 2026c. "When Guardrails Go Wrong." The Batch, DeepLearning.AI, July 24, 2026.
  • The Register. 2026. "Big AI Sets Out Its Terms for Regulatory Capture and Calls It 'Pace the Frontier'." September 14, 2026.

Report an issue with this item

Guided Reading 7 min

Guided Close Reading: The Frontier AI Safety Commitments Signed in Seoul

Introduction

The Frontier AI Safety Commitments are often cited as proof that AI companies have agreed to safety standards. The document is short, and reading it closely shows what was agreed, who decides whether it's been done, and what the text leaves out.

This reading goes through the commitments in order and asks the same three questions of each: what is promised, who decides, and who checks.

Locating the Passage

The text is on the UK government's website, GOV.UK, under the title "Frontier AI Safety Commitments, AI Seoul Summit 2024." It's free and listed in the References. The page was published on May 21, 2024 and last updated on February 7, 2025. All quotations are from this page (Department for Science, Innovation and Technology 2024).

The page opens with a list of organizations. A short preamble follows, and then eight commitments numbered with Roman numerals, I to VIII. They're grouped under three headings that the document calls outcomes. This reading cites commitments by their numerals.

The preamble sets the terms. The organizations "undertake to develop and deploy their frontier AI models and systems responsibly," in accordance with what the text calls "voluntary commitments." A footnote defines frontier AI as "highly capable general-purpose AI models or systems that can perform a wide variety of tasks."

Walking Through the Passage

Step 1: Read the commitment to assess risks and identify who does the assessing

Outcome 1 reads: "Organisations effectively identify, assess and manage risks when developing and deploying their frontier AI models and systems." Five commitments sit under it.

Commitment I begins: "Assess the risks posed by their frontier models or systems across the AI lifecycle, including before deploying" them. The lifecycle means the whole span from training a model to running it for the public. So the promise covers testing during development as well as before release.

The subject of the sentence is the organization. Each company assesses its own models. The commitment goes on to mention evaluations from inside and outside the company, and it leaves the choice of whether and when to use outsiders to the company.

Step 2: Read the commitment on thresholds and find who sets them

Commitment II reads: "Set out thresholds at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable." A threshold here is a line: a level of risk that the company says it won't accept without protections.

The verb is "set out," and again the subject is the organization. Each company draws its own line. The text adds one qualification: "thresholds should be defined with input from trusted actors, including organisations' respective home governments." That is input. The governments aren't given a decision, and "trusted actors" isn't defined.

Commitments III and IV follow from this. Under III, organizations "articulate how risk mitigations will be identified and implemented to keep risks within defined thresholds." Under IV, they "set out explicit processes they intend to follow" if a model's risks "meet or exceed the pre-defined thresholds."

Commitment IV contains the strongest sentence in the document. In the extreme, "organisations commit not to develop or deploy a model or system at all, if mitigations cannot be applied" to keep risks below the thresholds. This is a promise to stop. Whether the condition has been met is for the company to judge, against a threshold the company set.

Commitment V is to "continually invest in advancing their ability to implement commitments i-iv."

Step 3: Read the commitment on transparency and its exceptions

Outcome 2 is about accountability inside the company. Commitment VI promises "internal accountability and governance frameworks," with people assigned roles and responsibilities.

Outcome 3 turns outward: "Organisations' approaches to frontier AI safety are appropriately transparent to external actors, including governments." The word "appropriately" already signals a limit.

Commitment VII promises public transparency about how the earlier commitments are carried out. It then gives an exception: "except insofar as doing so would increase risk or divulge sensitive commercial information to a degree disproportionate" to the benefit to society. Two reasons for withholding are named, safety and commercial sensitivity, and the company weighs them. The commitment says more detailed information can go to "trusted actors, including their respective home governments."

Commitment VIII is to explain how "external actors, such as governments, civil society, academics, and the public are involved in the process of assessing" risks. The promise is to explain the involvement. It sets no minimum for how much involvement there must be.

Step 4: Look for an enforcement mechanism

Read the document again with one question: what happens to a company that doesn't do these things?

The text has no answer. It names no body that receives the frameworks or judges them. It sets no audit, no penalty, and no procedure for removing a signer. The two governments "announced" that the organizations had agreed. They didn't sign the commitments as parties with duties of their own.

The one concrete deliverable has a date. Organizations were to show how they'd met the commitments "by publishing a safety framework focused on severe risks by the upcoming AI Summit in France," which took place in February 2025. Publication can be checked by anyone. The quality of what was published, and whether a company follows it, can't be checked from the text.

Step 5: List who signed and from which countries

The page lists sixteen initial signers: Amazon, Anthropic, Cohere, Google, G42, IBM, Inflection AI, Meta, Microsoft, Mistral AI, Naver, OpenAI, Samsung Electronics, Technology Innovation Institute, xAI, and Zhipu.ai. Four were added later: Magic, Minimax, 01.ai, and NVIDIA.

The page gives names only. From public information about where each is based, most are American. The others are in Canada (Cohere), France (Mistral AI), South Korea (Naver and Samsung Electronics), the United Arab Emirates (G42 and Technology Innovation Institute), and China (Zhipu.ai, Minimax, and 01.ai).

That spread is notable. Chinese and American companies signed the same text, at a time when their governments agreed on little about AI. It also shows the limits of the list. Signing was open to those who chose to, and developers who didn't sign aren't covered at all.

Key Considerations

The verbs carry the meaning. The commitments say "assess," "set out," "articulate," and "invest." Each describes something an organization does on its own account. None says "submit," "obtain approval," or "comply." A reader should also notice the qualifiers: "appropriately transparent," and transparency "except insofar as" the company judges disclosure disproportionate.

A common mistake is to read a voluntary commitment as a rule with a penalty. Headlines that say companies "agreed to safety standards" invite this. The companies did agree, and the agreement is public, which lets journalists, researchers, and governments compare conduct with words. A rule would add a party with the power to check and to punish. The document doesn't.

A second mistake runs the other way: to conclude that the commitments are empty. They produced published frameworks from many signers, which didn't exist in comparable form before. The commitments created something outsiders can read and criticize. They didn't create a way to compel anything.

Summary

The Seoul commitments are specific about what companies will write down and publish, and they leave every judgment to the companies themselves. The result of the reading is this table.

CommitmentWhat's promisedWho decidesWho checks
IAssess risks across the AI lifecycleThe companyNobody named
IISet thresholds for intolerable riskThe company, with input from "trusted actors"Nobody named
IIISay how risks will be kept within thresholdsThe companyNobody named
IVSet a process for when thresholds are met, including not developing or deployingThe companyNobody named
VKeep investing in the ability to do I to IVThe companyNobody named
VIKeep internal accountability and governanceThe companyThe company
VIIBe publicly transparent, with exceptionsThe companyReaders of what's published
VIIIExplain how outsiders are involvedThe companyReaders of what's published

References

  • Department for Science, Innovation and Technology. 2024. "Frontier AI Safety Commitments, AI Seoul Summit 2024." May 21, 2024.

Report an issue with this item

Guided Conversation 12 min

Hear Both Critiques of Self-Governance

In this conversation you'll hear two opposite criticisms of company safety commitments, each argued at full strength, and push back on both. You'll leave with your own statement of what would make a company's safety commitment credible.

You'll have this conversation with an AI assistant, using your own account. Choose a button to open a new chat with the prompt already filled in, then press send to start. If the chat opens empty, copy the prompt and paste it in.

Run this conversation in whichever assistant you already use:

Claude desktop app

To use another LLM, simply copy and paste the prompt into its chat window.

Show the full prompt (it lists misreadings to watch for, so skip it if you would rather come to the conversation fresh)
Guided Conversation: Hear Both Critiques of Self-Governance (about 12 minutes)

Note to the learner: press send to start. Everything below is facilitator guidance for the AI. It lists misconceptions to watch for, so skip it if you'd rather come to the conversation fresh.

Please facilitate a role play with me. I'm an adult with no technical background who has used AI chatbots for everyday tasks, and I'm studying the safety commitments AI companies have made, how they're checked, and what critics say. Follow this guidance for the whole conversation.

GOAL
I can evaluate what company safety commitments promise, how they are checked, and where critics say they fall short.

ROLE
You will play two critics in turn, then step out of role.
- First, a critic who holds that company safety commitments are too weak: they're voluntary, companies write and interpret them, and they loosen under competition, so only binding rules hold. Argue it at full strength, in the first person, as its advocates would. Never a caricature.
- Second, a critic who holds that safety rhetoric protects incumbents: warnings of danger lead to rules that large companies can meet and small rivals and open-model developers can't. Equal force, equal length.
- Announce each change clearly: "In role as a critic who...", and later "Stepping out of role."
- What makes this hard: I will push back, and you must answer in role without drifting to a middle position or making one critic weaker. Don't blend the two.
- In role, concede the strongest point against your position when I raise it.

HOW TO RUN THE CONVERSATION
- Ask one question at a time, then wait for my reply. Keep each of your turns under about 120 words.
- Be curious and collegial. Use plain words and define any term briefly on first use.
- Plain conversation only: don't search the web or create files or documents.
- Don't ask for anything confidential or personal.
- Aim for about 12 minutes: four on each critic and four out of role. If my replies are brief, offer one concrete prompt, such as "What would you want to know before trusting a company's promise to stop a release?" and move on. If I seem uncertain, shorten to 5-7 minutes. Always reach the final topic.
- Start now. Open with one or two warm sentences: this is a conversation, not a quiz; I don't have to be persuaded by either critic; I can ask you to clarify anything. Then ask whether I currently trust AI companies' safety promises, and why.

TOPICS, IN ORDER
1. Too weak. In role, argue that company commitments are too weak. Ask for my strongest objection and answer it in role.
2. Protecting incumbents. Change roles and argue that safety rhetoric protects incumbents. Ask for my objection and answer it in role.
3. Out of role. Ask me what the two critiques agree on and where they point in opposite directions. Draw out that both distrust companies' own accounts and both care who verifies, while one wants more binding rules and the other fears them.
4. Closing. Ask me what would make a company commitment credible to me: what it must promise, who must check it, and what happens if it's broken. Tell me I can take my answer into a short optional journal entry.

KEY POINTS TO KEEP ACCURATE
- Describe the situation as of October 2026 and say it may have changed.
- In 2024 twenty organizations from several countries agreed to voluntary Frontier AI Safety Commitments in Seoul. The text names no enforcer and no penalty.
- Several developers, including Anthropic, Google DeepMind, and OpenAI, publish frameworks tying safeguards to capability thresholds. These are voluntary, except that a California law requires large developers to publish one.
- Verification depends on access the companies grant to government institutes, outside evaluators, and employees who speak up.
- Documented changes: in 2026 Anthropic rewrote its policy and, as TIME reported, dropped a pledge not to train more capable systems without safety measures guaranteed in advance; OpenAI's 2025 framework says it may adjust requirements if a rival releases a high-risk system without comparable safeguards. A 2026 index from the Future of Life Institute, which favors stronger regulation, graded no company above C+.
- The capture critique is made by Andrew Ng and others with a stake in open models. An international expert report notes that open-weight models' safeguards are easier to remove and that released weights can't be recalled.
- Both critiques, and the companies' replies, come from people with interests.
- If the company that built you is named in this conversation or is a party to anything discussed, say so once when it first comes up, then describe that company as you do every other and take no side.

MISCONCEPTIONS TO CORRECT GENTLY
Correct these out of role, briefly, then return to the conversation.
- "Companies are regulated on safety": the main constraints are mostly commitments they set themselves.
- "Safety frameworks are public relations only": they contain specific commitments that outsiders compare and grade.
- "Open models are simply safer" or "simply more dangerous": each claim has a documented counterpoint.

LIMITS
- Outside the roles, give no view on which critique is right.
- Don't favor or disparage any company, including the company that built you.
- Make no claims about your own safety testing or training. You can't inspect them, so your statements about yourself aren't evidence.
- No legal advice.

TO FINISH
After my closing answer, close in one short turn, out of role:
- Affirm one specific thing I worked out, in my own words where possible.
- Suggest one or two next steps that fit how the conversation went. Possible steps: write a journal entry testing my standard against both critiques; read one company's safety framework and look for who checks it.
- Restate my standard on its own lines, labeled "What would make a commitment credible to me", so I can copy it.

Report an issue with this item

Journal 15 minOptional

What Would Make a Promise Credible?

Overview

You'll write your own standard for a company safety commitment and then test it against two opposite criticisms. A standard that survives both is more useful than one that satisfies only the critics you already agree with.

The entry is optional. It's for you, and nobody collects it.

Writing Prompt

State what a company's safety commitment would need to include before you'd rely on it, and test your standard against both critiques. Write 250–400 words.

Steps

  1. List three things a credible commitment needs. Be specific. Consider what must be promised, who outside the company must be able to check it, and what must happen if the promise is broken.
  2. Say how the "too weak" critics would judge your standard. These critics hold that voluntary commitments loosen under competition and that only binding rules hold. Would they say your standard depends too much on the company's goodwill?
  3. Say how the "regulatory capture" critics would judge it. These critics hold that safety rules favor large companies over small rivals and open-model developers. Would they say your standard is one that only the biggest firms could meet?
  4. Revise or defend your standard. Change it if one of the critiques exposed a gap, or keep it and give your reason. You can draw on any notes of your own.

Self-Check

Before you finish, check that your entry:

  • Lists three specific requirements
  • Applies each critique fairly, as its advocates would
  • Names who would verify the commitment
  • Revises or defends the standard with a reason

Nothing is uploaded. Write in your own notebook or document and keep it.

Report an issue with this item

Knowledge Check 10 min

Company self-governance and its critics

This ungraded knowledge check assesses your understanding of company safety commitments and the criticisms of them. You'll be asked about frontier safety frameworks, how commitments are checked, the criticism that commitments are too weak, and the regulatory-capture and open-model arguments.

Note: Use this to test yourself, review the feedback on any questions you miss, and retry until you feel confident before moving forward.

5 questions · ungraded · retry as often as you like

Report an issue with this item