THE SHORT ANSWER
Keep meaningful human involvement when work is consequential, ambiguous, sensitive, exceptional or legally accountable. Decide whether the system should automate, assist or prepare an action for approval—and ensure the reviewer has time, context and authority to intervene.
AUTOMATE vs ASSIST vs REQUIRE APPROVAL
| Mode | Use when | Example |
|---|---|---|
| Automate | Rules are stable and failure is low-impact/recoverable | Route a complete request |
| Assist | A person benefits from a draft or classification | Suggest a support response |
| Require approval | Action is consequential or hard to reverse | Release payment or send sensitive communication |
Keep accountable review where it matters
- Financial commitments and destructive actions
- Customer complaints and vulnerable circumstances
- Hiring and employment decisions
- Legal or compliance judgments
- Sensitive external communication
- Strategic trade-offs
- Unusual exceptions outside tested policy
A nominal reviewer is not enough
Human oversight fails when the reviewer lacks context, cannot challenge the system or simply approves a queue under time pressure. Show inputs, uncertainty and consequences; make rejection and escalation usable.
The deeper human approval guide covers agent controls.
Evidence & context: NIST
Apply the idea to one real workflow
Choose one current workflow. Record the present outcome, the proposed change, the accountable owner, the most important exception or failure, and one before-and-after measure. Test the smallest safe version before expanding it.
Sources & further reading
- Generative Artificial Intelligence Profile (NIST AI 600-1)
NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.
- Demystifying evals for AI agents
Anthropic. A provider's engineering guidance on multi-turn agent evaluation, checked 13 September 2026. Examples inform evaluation design but do not establish universal pass thresholds.
- Authorization Cheat Sheet
OWASP Foundation. Security guidance emphasizing least privilege, deny-by-default behavior and authorization checks on every request. Implementation details depend on the application's threat model.
Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.
A correction, a counterexample or an experience worth sharing?
Join the conversation ↗