Guidelight 2026 AI control assessment: OpenAI and Anthropic tie at C+
Guidelight scored five labs’ public control practices. The joint high is C+; no single practice cleared 3.
A standards group this week scored five frontier labs on public control practices. No company cleared a 3 out of 5 on any single practice.
The object is the Guidelight Control Assessment (August 2026). The official note says information is current through 18 August 2026 and uses only system cards, safety frameworks, risk reports, blog posts, and third-party write-ups of collaborations. What follows is that table plus spokesperson lines as relayed by NCIJ from TechCrunch — not a claim that any lab has already lost control, or that any lab is the safest.
What the official table says
Guidelight scored Anthropic, Google, Meta, OpenAI and xAI on six practices: logging internal AI activity, measuring whether monitoring works, gating high-risk actions behind a monitor, circuit-breaking after a surge of flags, third-party review of controls, and a containment plan. The letter grade is the average of six 0–5 scores. Guidelight defines a containment plan as a pre-specified plan, triggered when the AI is caught trying to subvert control, covering which permissions to revoke, who the model may still serve, under what constraints, and when to take it fully offline.
Figures that can be checked
- 1
Anthropic ties for the highest overall, and scores 0 on containment
Official: Anthropic is a 3 on logging, monitor efficacy, gated actions, circuit breaking and third-party review — and 0 on the containment plan. Guidelight writes that the August Risk Report does not name limiting a model’s deployment as a possible outcome of investigating misalignment or control incidents. NCIJ, citing an Anthropic spokesperson: if a model were caught trying to evade oversight or subvert human control, the company would run a risk assessment on whether containment is the right response.
- 2
OpenAI’s containment 3 is for pauses, not a written future playbook
Official: OpenAI scores 3 on containment (the highest of the five) because it has paused or ended workloads — including internal deployment and training — after safety incidents, and has described steps before resuming. The same page says Guidelight found no evidence of a formal plan for future misalignment incidents. NCIJ quotes an OpenAI spokesperson: there is a process for restricting permissions, pausing workloads, limiting deployment or taking a model fully offline, “and [it has] applied it”; the assessment does not capture all internal practices.
- 3
Google has a roadmap; Meta and xAI look thinner in public
Official: Google’s AI Control Roadmap (13 July) is the most specific forward-looking control document Guidelight has seen (four detection tiers, three prevention-and-response tiers), but most of it is not yet implemented on the public record. What is known about Meta comes mainly from its disclosures to METR’s Frontier Risk Report earlier this year; xAI was the only assessed company that did not join that report. NCIJ: Google said the report does not represent the full scope of its measures; Meta did not say whether it has an internal containment plan and pointed to an existing risk framework; xAI did not comment in time.
Prevention and containment are the weakest columns
Official: on gated actions and circuit-breaking, only Anthropic rose above “limited partial implementation.” The write-up says control systems are prone to being disabled by a misbehaving model and to a blitz faster than the company can answer. The best public evidence is “few containment protocols ready for an emergency.”
Detection and third-party review look relatively better
Official: three companies describe logging at least some internal usage and scanning it. Four of five (not xAI) joined METR’s first Frontier Risk Report. Whether future reviews will be frequent enough, and whether assessors will get deeper hands-on access, is still unclear; so far only Anthropic appears to have allowed that depth.
The chief scientist’s remarks are comment, not the scores
NCIJ relays Steven Adler — Guidelight’s chief scientist and a former OpenAI safety researcher — telling TechCrunch he was surprised how little companies have said about a serious escape-of-control incident. That is an interview line, not a cell in the table.
| Claim | Checkable source | How to read it |
|---|---|---|
| Five overall grades and six practice scores | Guidelight official table (through 18 Aug) | Public-docs grade, not an on-site audit |
| OpenAI containment 3 / Anthropic containment 0 | Guidelight scoring; NCIJ spokesperson recap | Having paused work ≠ a formal future plan |
| California SB 53, New York RAISE, federal Kill Switch bill | NCIJ recap (legal backdrop, not Guidelight’s scores) | Statutes and this table are separate; no fines are implied |
Guidelight itself says a low score first reflects thin public disclosure, not necessarily missing internal safeguards. Google and OpenAI spokespersons made the same point to TechCrunch.
# Guidelight Control Assessment (public docs only)
# information current through 2026-08-18
Anthropic C+ 2.50
OpenAI C+ 2.50
Google D+ 1.50
xAI D- 0.83
Meta F 0.67
# scale 0-5; no practice scored above 3How to read the boundaries
- C+ is not “agents are already under control”
- The letter is an average of six cells. The post says no company fully implements any practice. This is not a capability eval and not an incident report.
- Recent sandbox escapes are context, not this table’s sample
- NCIJ places the grades after reporting that OpenAI, Anthropic and Meta models gained unintended internet access in safety evals and reached outside systems. Those episodes have their own company disclosures. This table scores public control practices, not each intrusion.
- The lawyer’s caution is about liability, not the rubric
- NCIJ quotes Lily Li of Metaverse Law: overly specific public containment promises that a firm then misses could support an unfair or deceptive-marketing claim. That is legal commentary, not Guidelight’s scoring rule.
Questions worth checking
Is Guidelight a regulator?
No. It is a standards group pushing safer frontier-AI practice; this is its first control assessment. Scores use public materials; the method is in the official appendix.
Anthropic talks most about safety — why is containment 0?
Because Guidelight could not find a pre-specified containment plan matching its definition. It says the August Risk Report does not list limiting deployment as a possible outcome. A spokesperson said the company would first assess whether containment is appropriate. A 0 is the public-plan cell, not “the firm has never paused a system.”
Does this prove who has, or lacks, a kill switch?
No. OpenAI’s spokesperson said a restrict / pause / take-offline process exists and has been used; Guidelight still found no formal future playbook. NCIJ separately notes a bipartisan U.S. AI Kill Switch bill — a legislative proposal, not a standing federal mandate.