Look for bounded, repeatable work
- Classification into a stable set of labels.
- Extraction into a validated schema.
- Rewriting within clear tone and factual constraints.
- Routing requests whose ambiguous cases can escalate.
- High-volume drafting where a person already reviews the result.
A small model is not automatically safe. Suitability comes from evaluation, controls and consequence—not the task's familiar label.
Recognise escalation signals
| Signal | Possible response |
|---|---|
| Missing or conflicting evidence | Retrieve again or ask a person |
| Input outside the tested distribution | Reject or route to a broader model |
| Invalid structure after one repair | Fail clearly rather than loop |
| High-consequence decision | Require authorised human judgment |
| Complex synthesis or nuanced instruction | Test a more capable model |
Pilot with a shadow comparison
Run the smaller option beside the current process without letting it take consequential action. Compare accepted results by category, not just overall. Inspect disagreements and severe errors. Then introduce it gradually with logging and a fallback.
A representative evaluation also prevents an easy majority class from hiding poor performance on rare but important cases.
Evidence & context: Microsoft Learn · NIST
Keep the boundary current
The goal is not to maximise use of smaller models. It is to allocate capability deliberately and preserve a clear route for cases the default cannot handle.
Sources & further reading
- Models
OpenAI Developers. Official model-selection documentation, checked 13 September 2026. Product names, capabilities and prices can change; the collection uses the durable principle of matching capability to a task rather than prescribing a current model.
- Evaluate a model router
Microsoft Learn. Official guidance for evaluating routing across representative workloads using quality, cost, latency and policy criteria. It is not evidence that routing always improves results.
- Generative Artificial Intelligence Profile (NIST AI 600-1)
NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.
Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.
A correction, a counterexample or an experience worth sharing?
Join the conversation ↗