THE SHORT ANSWER
AI quality control begins with an explicit task, acceptable output and unacceptable failure. Test representative and difficult cases, measure error and human-review performance, establish fallback behavior, control changes and monitor production rather than relying on a demonstration.
Define quality before selecting a metric
| Dimension | Question |
|---|---|
| Correctness | Is the output factually or operationally right? |
| Completeness | Does it omit material requirements? |
| Consistency | Does similar input receive appropriately similar treatment? |
| Robustness | What happens with unusual, adversarial or poor input? |
| Usability | Can the person understand and act appropriately? |
Create a failure catalogue
Record hallucination, unsupported inference, stale information, formatting failure, unsafe action, excessive refusal and human-review misses. Weight failures by consequence rather than reporting one average alone.
Put quality checks inside the workflow
- Retrieve from approved current sources where appropriate.
- Validate structured outputs before use.
- Require evidence for consequential claims.
- Route uncertainty and exceptions to people.
- Fail safely when dependencies are unavailable.
- Version prompts, models and evaluation sets.
A passing system can regress
Model, prompt, data, tool, policy or user changes can alter behavior. Define which changes trigger retesting and compare production signals with the approved baseline.
Use the existing guides on hallucinations and verification for individual output checks.
Evidence & context: NIST · National Institute of Standards and Technology
Sources & further reading
- Generative Artificial Intelligence Profile (NIST AI 600-1)
NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.
- Artificial Intelligence Risk Management Framework (AI RMF 1.0)
National Institute of Standards and Technology. Voluntary, rights-preserving guidance organized around GOVERN, MAP, MEASURE and MANAGE. NIST was revising AI RMF 1.0 when checked on 28 September 2026, so organizations should verify the current version before formal adoption.
- NIST AI RMF Playbook
National Institute of Standards and Technology. Suggested actions for using AI RMF 1.0. It is voluntary, not a checklist or certification, and NIST states that it will be updated after the framework revision.
Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.
A correction, a counterexample or an experience worth sharing?
Join the conversation ↗