THE SHORT ANSWER
AI can draft code, explain unfamiliar components, suggest tests, help debug, document behavior, refactor and coordinate bounded coding tasks. Generated output can also be incorrect, insecure, outdated or poorly fitted to the wider system. A responsible workflow keeps requirements, review, tests and deployment controls around the model.
Use AI for bounded development tasks
| Task | Useful assistance | Required check |
|---|---|---|
| Prototyping | Draft a narrow interaction | Does it test the intended assumption? |
| Code generation | Produce an implementation candidate | Does it fit architecture and requirements? |
| Debugging | Suggest causes and diagnostic steps | Can the cause be reproduced and fixed? |
| Testing | Propose cases and test code | Do tests cover behavior rather than mirror code? |
| Refactoring | Restructure existing code | Did behavior, performance or security change? |
| Documentation | Summarize interfaces and decisions | Is it accurate and current? |
Generated code can be plausible and wrong
A model may invent an API, use an obsolete pattern, omit authorization, mishandle an edge case or create a dependency the team cannot support. Passing syntax is not evidence that the product behaves correctly.
The less a team understands the generated system, the harder it becomes to diagnose incidents, evaluate changes and control cost.
Keep a controlled software-delivery loop
- State the requirement and constraints.
- Give only the minimum necessary context and access.
- Generate or modify one bounded change.
- Review code and dependency choices.
- Run relevant tests and security checks.
- Inspect the product behavior.
- Commit a traceable change and monitor deployment.
A coding agent should receive scoped tools and permissions rather than unrestricted production access.
Measure accepted work, not generated volume
More code can increase review, maintenance and compute cost. Judge AI assistance by accepted outcomes, defect rate, time to recovery and the ongoing burden of the system.
Use Cost-efficient AI Workflows for the wider operating model. If the product calls an AI API, the API Cost Calculator can test usage assumptions.
Sources & further reading
- Generative Artificial Intelligence Profile (NIST AI 600-1)
NIST. Risk-management guidance, including confabulation. It does not establish a universal error rate.
- Secure Software Development Framework
NIST. Outcome-based secure-development guidance covering preparation, protection, secure production and vulnerability response. It is a framework, not a product-specific checklist.
- Quality assurance: testing your service regularly
GOV.UK Service Manual. Government guidance on usability, functional, performance, security, accessibility and continuous testing. It does not prescribe one universal test suite.
- Secrets Management Cheat Sheet
OWASP Foundation. Security guidance on secret creation, storage, distribution, rotation and revocation. It supports the principle that private credentials do not belong in public client code.
Examples and exercises are illustrative unless attributed to a source. No independent expert review is claimed.
A correction, a counterexample or an experience worth sharing?
Join the conversation ↗