September made AI auditing the default answer to a month of incidents. Frontier labs, their CEOs, and the White House all landed on the same substitute for a binding rule: someone independent should look at the model. That instinct is right. The paperwork so far is not an audit.
On September 29, 2026, President Trump and a group of AI executives signed the White House Accord on Super Intelligence after a lunch at the White House. Signers included Anthropic’s Dario Amodei, Google’s Sundar Pichai, Meta’s Mark Zuckerberg, OpenAI’s Greg Brockman, Nvidia’s Jensen Huang, and Elon Musk. Trump called it “morally binding.” It has no legal force. It commits companies to four layers of controls: robust internal controls, an internal team that checks those controls, an independent external auditor or evaluator, and a board committee that reviews what the internal and external reviewers report. The text says that over time these steps may be written into law. They are not law today.
Why audits showed up now
The month that produced the Accord was a run of agent incidents, the first reported intentional agentic attack filed with a regulator, government data breaches attributed to AI agents, and public warnings from current and former lab employees about catastrophic risk inside a decade. Amodei’s essay “Pacing the Frontier” asked every frontier lab to host embedded third-party evaluators, naming METR as the example, with access to training pipelines and not just the finished model. Altman and Musk backed the idea in public. The Accord is the political version of that ask: evaluation instead of a pause, and a board committee instead of a statute.
That trade is familiar. When governments cannot agree on a horizontal rule, they ask for assurance. Financial reporting did this. Drug manufacturing did this. Cybersecurity did this with SOC 2 and ISO 27001. The EU AI Act is still the only major binding conformity regime, and its core obligations do not land until December 2027. China already requires periodic compliance audits for large personal-information processors, plus algorithm filings for public-facing systems. The United States, at the federal level, just signed a one-page moral commitment.
An evaluator with a badge is not an audit
Three jobs are being collapsed into one word.
- Capability evaluation. Can the model do the dangerous thing? METR’s time-horizon tests, red teams, and pre-deployment evals live here. They answer a technical question. They do not certify a control environment.
- Assurance. Are the lab’s safety practices actually running? Incident reporting, training-pipeline review, weight security, refusal testing after fine-tunes. This is what Amodei means by an embedded evaluator.
- Audit. An opinion, against a stated standard, by someone who does not depend on the company for access, methodology, or the right to publish. This is the piece that does not exist yet for frontier models.
Brookings’ Elham Tabassi put the test in one line: who controls the methodology, who controls access, who funds the work, who controls publication. An evaluation is not credible because the person who ran it sits outside the building. Embedded access cuts both ways. It lets a reviewer see the training run. It also makes the reviewer a guest. Independence is a structure, not a job title. Financial auditors are banned from booking the revenue they later opine on. No equivalent rule exists for a lab that pays the evaluator, picks the tasks, and decides what the public sees.
Lina Khan’s point in the New York Times is the other half. These companies are already under consumer-protection law, sector rules, and a growing docket of lawsuits. Self-governance does not displace that. An accord that says “our guardrail is an auditor we hired” does not either.
The standards gap is the actual problem
Other safety-critical industries audit against a standard the auditor did not write. AI does not have that yet. NIST’s AI Risk Management Framework, ISO 42001, and the EU’s forthcoming harmonized standards are the closest things, and none of them is a frontier-model exam. Without a standard, the company funds the writing of its own test and hires someone to sit it. That is a report. It is not assurance a customer or a regulator can rely on.
The people doing the work have the same problem. Capability evals, control testing, and conformity assessment are different crafts. A red-teamer is not an IT auditor. An IT auditor is not a model evaluator. Certification and a scope limitation on the face of the report are how other fields keep those roles from blurring. AI assurance is early on both.
What is actually becoming law
The binding movement is in the states, and it is about the auditor, not the model.
- California. Newsom signed SB 813 and AB 1405 in September 2026. AB 1405 builds an AI auditor registry at the Government Operations Agency. From January 1, 2029, an unregistered person may not offer or conduct a covered AI audit in California, meaning an audit of controls needed for compliance with state law. SB 813 adds a voluntary, government-credentialed tier of independent verification organizations. An executive order later pulled the registry timeline forward.
- Companion chatbots. California’s SB 1119 is the first enacted bill aimed at that product category, and it pulls assessment duties in with it.
- Illinois and the federal mirror. Illinois SB 315 and the proposed Great American AI Act both reach for third-party assessment of catastrophic-risk models. The federal bill would license independent verification organizations and put those models on a six-month audit cycle. Neither is a finished national regime.
- Voluntary programs. Singapore and the United Kingdom already run assessment schemes companies opt into. Adoption is high because buyers ask for the artifact, not because a statute demands it.
Frontier commitments will land on customers whether or not those customers train a model. If a lab tells its board an external evaluator signed off, enterprise buyers will ask for the same letter. Procurement questionnaires will grow an AI-audit line the way they grew a SOC 2 line. Downstream deployers who fine-tune, wrap, or put an agent on a customer database inherit the incident, not the lab’s moral commitment.
What a real program looks like
A company that wants this to survive a regulator, a plaintiff, or a customer security review needs five things the Accord does not supply.
- A named standard. ISO 42001, NIST AI RMF, or a documented internal control set mapped to one of them. “We follow industry best practice” is not a criterion.
- A scope. Frontier weights, the agent scaffold, the tools it can call, and the logging around those tools are different systems. An opinion that does not say which one it covered is decorative.
- Independence in the contract. The evaluator keeps the methodology, the right to see failures, and the right to publish a summary. Payment does not include a veto.
- A board committee that can reject the report. The Accord’s fourth layer only works if the committee is not the same people who own the launch date.
- Evidence a customer can read. Incident log, eval results on the deployed system, and a suppression path when a control fails. A press quote from the lab is not that evidence.
Independent evaluation is the right tool
Independent evaluation is the right tool, and it is not new. Aviation, finance, and drugs already use it. What is new is treating a morally binding one-pager as the tool. The Accord points at internal controls, an outside evaluator, and a board committee. It does not say who writes the test, who pays the tester, or what happens when the tester fails the lab. California is further along, and it regulated the auditor before it regulated the model. Until a standard and a publication right exist, an AI audit is a private letter. Buyers should read it that way.