“A human will review the output” is not a control until the human has the time, evidence, and authority required to make the review meaningful.
On an architecture diagram, thoughtful oversight and an overloaded queue can look nearly identical. Both have a person-shaped icon between the machine and the decision. The diagram rarely shows the reviewer’s calendar, incentives, evidence interface, or ability to stop the work.
Use deliberately simple math. One agent-assisted process produces 240 items each day. Every item routes to one human reviewer. A meaningful review takes three minutes. That is twelve hours of review before breaks, context switching, hard cases, or the reviewer’s actual job. The numbers are illustrative, but they force the design to account for time. The capacity problem cannot be solved by making the approval button more visible.
The organization will resolve the mismatch somehow. It might add reviewers, reduce volume, narrow scope, automate a well-defined check, or let the queue accumulate. It might also allow review quality to deteriorate while preserving the visible act of approval. That last outcome is especially dangerous because the records still look reassuring. The process has a human signature. The human did not have the conditions required to make the signature mean what leadership thinks it means.
Agent capacity can grow faster than judgment capacity. Systems can generate, compare, draft, analyze, package, and produce at a pace that makes the human decision point look quaint. Then someone still has to decide which claims, risks, or actions they are willing to put their name on. Governance has to shape the flow of work before review becomes the place where impossible volume goes to look responsible.
A reviewable item should arrive with the decision already visible. What action is proposed? What purpose does it serve? What authority supports it? What changed? Which uncertainty matters? Where can the underlying evidence be inspected? “Please review” attached to a large generated document is not yet a well-formed request. The reviewer should not have to excavate the decision from a mountain of prose. Archaeology is already a profession.
Routing matters. Some checks can be deterministic. Some bounded actions can proceed under explicitly delegated authority. Some cases require specialist judgment. Some should stop because a prerequisite is missing. The classification itself needs testing; calling something low-risk does not make it so. But sending every item through the same expensive human judgment step creates a bottleneck that will change behavior somewhere else in the system.
The interface matters too. A concise summary can help, but only if the reviewer can inspect its basis and notice material omissions. If the same agent proposes the action and selects everything the reviewer sees, the design should account for that dependency. The person needs a usable way to challenge the account, not merely react to its confidence.
Measure review as actual work. Track queue age, observed review time, case mix, decisions changed by review, errors found afterward, and cases where reviewers could not obtain enough evidence. Compare a sample of approvals with deeper assessment. A fast approval may reflect a simple case or an excellent interface. It may also reflect a person who has learned that reading everything is impossible. Timing becomes evidence when paired with outcome and context.
The operating principle is straightforward: cadence governs demand, readiness governs admission, and sequencing governs execution. Put recurring work on a rhythm. Require materials before committing review. When unplanned work enters a full schedule, name what it displaces, use a deliberate reserve, or add capacity. There is no invisible fourth resource called “the team will somehow absorb it.” That resource usually turns out to be someone’s evening.
The oversight test fits in an ordinary operating review: show the volume, show the time, show the evidence the reviewer receives, and show a case where review changed the outcome. Then show how that continues to work as volume grows.
If the math requires a twelve-hour afternoon, the control has already failed.
