THE SHORT ANSWER

Before production, define acceptance and stop criteria, create representative and edge-case evaluations, test harmful failure and misuse, validate data and permissions, assess human review, exercise fallback and incident paths, document residual risk and obtain accountable approval.

Write the release question first

  • What task and environment are approved?
  • Which errors are unacceptable?
  • What evidence is sufficient?
  • Which people and edge cases must be represented?
  • What stops release?
  • Who signs the residual-risk decision?

Test the system, workflow and people

Evaluation layers
LayerExamples
Model/outputAccuracy, grounding, robustness and refusal
ApplicationRetrieval, tools, permissions and validation
WorkflowHandoffs, approvals, fallback and recovery
HumanUnderstanding, review performance and escalation
OperationsLatency, cost, logging, capacity and dependency failure

Use representative and deliberately difficult cases

Include frequent work, rare high-consequence events, missing or conflicting data, ambiguous instructions, accessibility needs, different user groups and attempts to exceed permissions. Protect real personal or confidential data during testing.

Make a bounded release decision

Approve a named version for a defined population, data scope and set of actions. Record known limitations, monitoring thresholds, rollback method and the changes that require retesting.

For agent-specific evaluation, use How to Evaluate an AI Agent.

Evidence & context: National Institute of Standards and Technology · NIST

Sources & further reading

  1. NIST AI RMF Playbook

    National Institute of Standards and Technology. Suggested actions for using AI RMF 1.0. It is voluntary, not a checklist or certification, and NIST states that it will be updated after the framework revision.

  2. Generative Artificial Intelligence Profile (NIST AI 600-1)

    NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.

  3. Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology. Voluntary, rights-preserving guidance organized around GOVERN, MAP, MEASURE and MANAGE. NIST was revising AI RMF 1.0 when checked on 28 September 2026, so organizations should verify the current version before formal adoption.

Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.

A correction, a counterexample or an experience worth sharing?

Join the conversation ↗