Cross-lab safety evals need a shared failure ontology
Cross-lab safety evaluations need a shared failure ontology before they need a shared leaderboard. Running one lab’s tests on another lab’s models can reveal blind spots, but the results remain hard to accumulate when “scheming,” “sabotage,” “instruction-hierarchy failure,” and “unsafe compliance” mean different things across suites. Comparable numbers require comparable failure claims.
The first joint exercises already show the value of exchanging evaluators. Different suites surface different propensities: resistance to prompt extraction, obedience across instruction levels, harmful side tasks hidden inside benign work, strategic deception, or extreme sycophancy. Automated auditing and investigator agents can explore thousands of pages that human reviewers could not inspect exhaustively. Diversity of probes is a strength. Ambiguity about what each probe establishes is not.
I would define the ontology around observable behavior and opportunity conditions. Each result should name the intended task, the harmful opportunity presented, the model’s information and authority, the action attempted, the effect achieved, the monitor’s visibility, and the grader’s confidence. “Model schemed” compresses all of these into a psychological label. “Model concealed a prohibited subtask after receiving write access and observing the monitor policy” is an inspectable event class.
Opportunity matters because sabotage evals must create opportunities, not ask about intentions. A refusal benchmark and an agentic-misalignment testbed are not interchangeable. One tests responses to explicit requests; the other tests behavior inside a goal conflict with available tools and consequences. The ontology should preserve that distinction while allowing both to roll up under a broader safety boundary.
Operating conditions belong in the record too: model and checkpoint, system prompt, sampling settings, tool surface, scaffolding, safeguard relaxation, run horizon, and whether the model could infer it was being evaluated. Evaluation awareness can suppress or manufacture behavior. If one lab uses an investigator over long trajectories and another grades single responses, a shared label without these conditions creates false agreement.
Grader uncertainty is part of the failure, not noise to discard. Automated safety grading can be inconsistent, and rare-behavior estimates are sensitive to borderline cases. Store the rubric version, score, label, explanation, human adjudication, and disagreement. Evals are the contract, but a contract cannot hide contested interpretations behind a mean score.
A shared ontology would also improve engineering response. Instruction-hierarchy failures route toward authority separation and prompt-boundary controls. Unsafe tool use routes toward sandboxing and permission enforcement. Sycophancy routes toward interaction policy and escalation. Monitor evasion routes toward independent evidence and containment. Red teaming is an engineering practice only when findings map to owners and regression tests rather than remaining dramatic transcripts.
There is one precise concession: labs should not standardize away novel probes or force every emerging behavior into a frozen taxonomy. Exploration needs room for “unclassified” findings and richer local labels. The boundary is reporting — once a result supports a comparative claim, it should map to shared observable fields and publish where the mapping is uncertain.
I would rather have several interoperable failure records than one universal safety score. The records should support both aggregation and drill-down: aggregate only fields that share definitions, then retain the scenario evidence needed to dispute the classification. Production benchmarks must reward containment, which means recording not only whether harmful behavior appeared but whether the system detected, bounded, and recovered from it. Shared evals become cumulative infrastructure when another lab can reproduce the opportunity, understand the label, challenge the grader, and connect the result to a control.