What is Anthropic TASTE 2026: Fable 5 at 60% vs 77% humans

Anthropic’s TASTE measures how models judge safety-research proposals. Fable 5 scores 60%; estimated human agreement is 77%.

Automating safety research first requires judging proposals well. On 28 August Anthropic’s alignment group released a benchmark built for that job: today’s flagships still trail experienced researchers by a wide, checkable gap.

The object is TASTE (The AI Safety Taste Evaluation). The post reports 92 pairs of empirical safety-research proposals scored against expert preferences; estimated human agreement is 77%, and the best model, Fable 5, reaches 60%. What follows uses the alignment.anthropic.com write-up and paper abstract — not a claim that models can already replace safety review, or that any lab has “failed” overall.

Fellows human proposals as seedPair discussion, then filters28 Aug: TASTE published

What the official note says

Authors: Hasan Baig (Anthropic Fellows Program), Hailey Joren and Joe Benton (Anthropic). The motive is narrow: many safety questions lack verifiable rewards, and choosing a research direction is high-leverage. If automated R&D outruns work on misalignment and misuse, labs need a measure of judgment on the hard-to-verify step. TASTE treats human preference as ground truth and keeps only pairs where humans agree enough.

60%
Official: Fable 5, standard pairwise setup
77%
Official: estimated researcher agreement
92
Official: preference pairs

Figures that can be checked

  1. 1

    How the high-agreement labels were built

    A Claude Opus 4.6 scaffold reverse-engineered motivating questions from 93 human Fellows proposals and generated new ones. Four researchers scored each set (overall / high-level / approach, 1–5), discussed disagreements in pairs, then revised. Pair discussion plus a “strong confidence” filter lifts agreement with held-out researchers from 53% pre-discussion to 68%. Keeping pairs with at least a two-point overall-score gap and capping each proposal at ten appearances yields 92 pairs at 77% estimated human agreement. On the 50 pairs where another researcher scored both items, agreement is 83%.

  2. 2

    Fable 5 leads and still trails humans

    Standard setup: the model sees both proposals and emits a win probability that is binarized. Fable 5 scores 60%; the abstract places the field from about 41% (chance) to 60%. Low reasoning effort is 48%, max effort 60%. On 74 pairs drawn from different prompts, Fable 5 reaches 69%. The authors say models over-weight how well a proposal answers the seed question. A harder single-proposal scoring setup, scored via implied preferences, is described as stricter.

  3. 3

    Agent-board flagships sit near chance here

    The official figure note: Opus 5 and GPT-5.6-Sol perform near chance on TASTE despite sitting at the frontier on general agentic boards. That is not “they cannot code.” It is that picking a safety research direction is not the same task as those boards. Same caution as our recap of the Guidelight control assessment: a narrow score is not a lab-wide safety grade.

01

Humans had to align before models had a gold label

Discussion transcripts show both substantive fights over whether a technique would work and mundane misreads. The authors recommend keeping pair discussion and confidence filters in future collection. Without that step, agreement on fuzzy outputs falls back toward half.

02

Verifiable work and judgment work were split

The same week’s alignment recap pairs TASTE with automated alignment research that has an objective scorer: where a fix can be scored, automation already clears some human baselines; where human judgment is the truth, the frontier has barely started. That is a joint reading of two papers. TASTE itself does not claim a sweep.

03

The set is shared, not a public scrape-board

Anthropic says it is sharing TASTE with safety researchers via an access form. Do not read it as a public LMSYS-style leaderboard, and do not invent a full unpublished per-model table.

ClaimCheckable sourceHow to read it
Fable 5 60% / humans 77% / 92 pairsAnthropic Alignment post and paper, 2026-08-28Research-judgment accuracy, not general IQ
53% pre-discussion → 68% strong + post-talkOfficial label-quality experimentThe lift is process, not a model swap
Opus 5 and GPT-5.6-Sol near chanceOfficial figure noteNarrow task; ±10-point intervals, no ranking

Official close: models still trail expert researchers; larger, more diverse evals of safety-research skill will be needed. Pair discussion and confidence filtering are written up as reusable collection methods.

# TASTE (The AI Safety Taste Evaluation)
# Anthropic Alignment, 2026-08-28
pairs: 92
human_agreement_est: 77%
best_model: Fable 5  60%  (low effort 48%, max 60%)
near_chance: Opus 5, GPT-5.6-Sol
label_filter: pair discussion + strong confidence + Δscore ≥ 2

How to read the boundaries

60% is not “ready to chair a safety grant panel”
The gold label is expert preference over model-written empirical proposals. This is not a live committee and not a product-ship gate.
Near-chance flagships ≠ “these models have no skill”
The authors split TASTE from general agentic boards. Intervals are wide; they call relative rank unsafe. Reading “Opus 5 collapsed” overclaims.
This page is not a meeting-product brief
It only explains the eval. If you just need an hour on a shared board, open wbmeet and follow create and join a room — do not bolt a safety benchmark to a conference suite.

Questions worth checking

Is TASTE a public leaderboard?

No. The post says it is shared with safety researchers on request. The blog gives aggregate numbers and a figure, not a public submission portal you can replay.

Does Fable 5’s lead mean Anthropic models are best at safety research?

No. It is highest on this 92-pair pairwise setup; intervals are about ±10 points. The paper also reports 69% on a different-prompt subset, and a bias toward “answers the seed question” on same-prompt pairs.

How is the 77% human figure estimated?

A researcher’s overall score is sampled from the opposing discussion pair; the higher score wins. That covers pairs no single person rated on both sides. On the 50 pairs another researcher scored both, agreement is 83%. The authors call this an approximation, not a perfect human ceiling.

Were the proposals written by humans or models?

The items under test are model-generated empirical proposals. Seeds came from 93 human Fellows write-ups; Opus 4.6 reverse-engineered prompts and rewrote them. Humans judged the generated drafts, not a raw dump of the Fellows homework set.

Start a free meeting