Look for bounded, repeatable work

  • Classification into a stable set of labels.
  • Extraction into a validated schema.
  • Rewriting within clear tone and factual constraints.
  • Routing requests whose ambiguous cases can escalate.
  • High-volume drafting where a person already reviews the result.

A small model is not automatically safe. Suitability comes from evaluation, controls and consequence—not the task's familiar label.

Recognise escalation signals

Signals that the default path needs help
SignalPossible response
Missing or conflicting evidenceRetrieve again or ask a person
Input outside the tested distributionReject or route to a broader model
Invalid structure after one repairFail clearly rather than loop
High-consequence decisionRequire authorised human judgment
Complex synthesis or nuanced instructionTest a more capable model

Pilot with a shadow comparison

Run the smaller option beside the current process without letting it take consequential action. Compare accepted results by category, not just overall. Inspect disagreements and severe errors. Then introduce it gradually with logging and a fallback.

A representative evaluation also prevents an easy majority class from hiding poor performance on rare but important cases.

Evidence & context: Microsoft Learn · NIST

Keep the boundary current

The goal is not to maximise use of smaller models. It is to allocate capability deliberately and preserve a clear route for cases the default cannot handle.

Sources & further reading

  1. Models

    OpenAI Developers. Official model-selection documentation, checked 13 September 2026. Product names, capabilities and prices can change; the collection uses the durable principle of matching capability to a task rather than prescribing a current model.

  2. Evaluate a model router

    Microsoft Learn. Official guidance for evaluating routing across representative workloads using quality, cost, latency and policy criteria. It is not evidence that routing always improves results.

  3. Generative Artificial Intelligence Profile (NIST AI 600-1)

    NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.

Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.

A correction, a counterexample or an experience worth sharing?

Join the conversation ↗