Part of the series Formation Governance

Oversight is easiest to support before it changes a decision.

At the beginning of an evaluation arrangement, everyone can agree that transparency matters. The charter is being drafted. The right people have joined the meeting. Someone may even be getting a badge. The harder test arrives later, when the evaluator finds something material, the release matters, the team is tired, and the difference between “we should understand this better” and “we can proceed” has a price attached to it.

That is the meeting an oversight model has to survive.

In his September essay, Dario Amodei argues for embedding independent evaluators inside frontier AI companies, with continuing access to systems and people and rights to publish findings. He draws on a banking-supervision analogy, which is useful because supervision is not just a report. It is an operating relationship with access, escalation, independence, and consequences. We Must Pace the Frontier

The practical questions are not glamorous, but they decide whether the arrangement works. Access to what? In what format? With which logs, interviews, system records, and authority to ask follow-up questions? How does the evaluator know whether the evidence set is complete enough for the claim being made? What happens when the team believes the evaluator has misunderstood a test result? Who owns the decision if the disagreement remains unresolved?

There are ordinary ways to starve an evaluator who technically has access. The relevant data can be available but unusable. Logs can omit the behavior that matters. Every question can require an introduction to someone who is busy, traveling, or not quite sure who owns the current version. A badge opens doors. It does remarkably little about naming conventions.

METR has already named several important conditions for serious investigation: model and transcript access, employee interviews, adequate resources, communication with oversight bodies, and disclosure of how redactions affect conclusions. That gives the field something concrete to build on. METR’s investigation framework

The next step is to make those conditions visible with the finding itself. A useful report should show what investigators requested, what they received, what remained unavailable, and what each gap prevented them from determining. “No evidence of a problem” means one thing after direct access to the relevant behavior and another thing when the behavior was never logged. Sensitive evidence may need restricted handling; the public account should still explain how the restriction affected the conclusion.

The disagreement path matters just as much. If an evaluator identifies a concerning behavior and the lab believes it is an artifact of the test, that may be true. The right response is a documented competing explanation and a way to distinguish between them. Preserve the original observation, the alternative account, the evidence that would resolve the question, and the decision taken while uncertainty remains. A finding should be able to change because the evidence changed. It should also be possible to tell when the language changed because the meeting got uncomfortable.

Consequences need the same clarity. An embedded evaluator does not need unilateral control over company operations to matter. A serious finding does need a route to someone who can accept the risk, require a remedy, restrict an action, or explain why proceeding is justified. That decision needs an owner and a record. Otherwise the organization can comply with evaluation indefinitely while leaving the underlying condition untouched. The report becomes another artifact the system knows how to produce.

The most useful preflight exercise is simple: run a bounded disagreement before the stakes are existential. Give the evaluator an incomplete evidence set, a disputed finding, and a real escalation path. Watch how long it takes to reach the right people. Watch whether uncertainty survives the executive summary. Watch who can request more work, who pays for it, and whether the decision-maker receives the evaluator’s actual conclusion.

The welcome meeting can tell us that everyone supports oversight. The first consequential disagreement tells us whether we built it.

If the finding cannot leave the building, the evaluator never really got in.