AI coding agents have moved from autocomplete to doing whole tasks: reading a codebase, running commands, editing files and checking their own work. Adoption grew quickly through 2026, and so did a familiar complaint in developer surveys: the output is often almost right, and debugging almost-right code takes time.
Most of that gap comes from how the work is set up, not from the model.
Give the agent a way to check itself
The single biggest improvement is a fast, reliable feedback loop.
- A test command that runs in seconds, not minutes.
- A type checker and linter the agent can run.
- A clear instruction: "run the tests and type check before you say you are done".
An agent that can verify its work iterates toward correct code. An agent that cannot is guessing.
Write a project instructions file
Most agents read a project-level file (for example CLAUDE.md or AGENTS.md). Keep it short and specific:
# Project notes
- Package manager: pnpm. Never use npm.
- Run `pnpm test` and `pnpm typecheck` before finishing.
- API handlers live in src/server/routes. Follow the pattern in users.ts.
- Do not edit generated files in src/gen.Rules that prevent repeated mistakes are worth more than a description of the architecture.
Scope tasks like you would for a new teammate
- One outcome per task. "Add pagination to the orders endpoint" beats "improve the orders feature".
- Point to examples. "Follow the pattern in
users.ts" saves the agent from inventing a new one. - State constraints. Performance limits, libraries you do not want, files that must not change.
Plan first for bigger changes
For anything touching more than a few files, ask for a plan before code. Review the plan, correct it, then let the agent implement. Fixing a plan is cheaper than reviewing a wrong 600-line diff.
Review like it matters
Agent output deserves the same review as human output, plus a few extra checks:
- Did it change tests to make them pass instead of fixing the code?
- Did it add dependencies you did not ask for?
- Are there hard-coded secrets, disabled checks or broad
try/catchblocks hiding errors? - Does the diff contain unrelated "cleanups"?
Protect your environment
Agents run commands. Treat that seriously.
- Use permission modes or allowlists for commands that write, delete or reach the network.
- Keep production credentials out of the development environment.
- Be careful with agents reading untrusted content (issues, web pages, dependency READMEs), which can contain prompt injection.
Key takeaways
- A fast test and type-check loop matters more than prompt wording.
- A short project instructions file prevents repeated mistakes.
- Scope tasks tightly and plan before large changes.
- Review agent diffs carefully and limit what commands can do.