Prompt engineering is about writing one good instruction. Context engineering is about everything else the model sees: system instructions, tool definitions, retrieved documents, previous turns and tool output. Once an agent runs for more than a few steps, that "everything else" dominates the result.

The core problem is simple. Models get measurably worse as the context grows, long before the window is full. More tokens is not more intelligence. The job is to decide which tokens earn their place at each step.

The four moves: write, select, compress, isolate

Most practical techniques fall into four groups.

  • Write – persist information outside the window. Notes files, a scratchpad, a task list, a database row. The agent re-reads what it needs instead of carrying everything.
  • Select – pull in only what is relevant now. Retrieval, file search, grep, a targeted query. Precision beats volume.
  • Compress – summarise what you must keep. Replace a 4,000-line log with the three lines that matter, or old turns with a short running summary.
  • Isolate – give a sub-task its own clean context. A sub-agent that researches one question returns a paragraph, not its entire browsing history.

Start with the system prompt

The system prompt is read on every step, so it is the most expensive real estate you own.

  • State the goal, the constraints and what "done" looks like.
  • Prefer concrete rules over adjectives. "Run the test suite before reporting success" beats "be thorough".
  • Remove anything the model would do anyway. Every sentence should change behaviour.

Treat tool definitions as context

Tool names, descriptions and schemas are tokens too, and the model reasons over them.

  • Keep the tool list small for each task. Fifty tools make selection harder than five.
  • Write descriptions that say when to use a tool, not only what it does.
  • Return compact, structured results. A tool that dumps 200 KB of JSON poisons every later step.
typescript
// Return what the agent needs, not everything the API gave you
return {
  status: run.status,
  failedTests: run.failures.slice(0, 5).map(f => ({ name: f.name, message: f.message })),
  totalFailures: run.failures.length,
};

Manage long-running history

For multi-step agents, history grows fast. Useful patterns:

  1. Rolling summaries – every N turns, replace older turns with a summary of decisions and open questions.
  2. Clear stale tool output – once a file has been edited, the old read of it is misleading as well as wasteful.
  3. Checkpoints on disk – write progress to a PROGRESS.md or task list, so a fresh context can resume.

Retrieval: precision over recall

When you retrieve documents, fewer and better beats more.

  • Chunk by meaning (sections, functions), not by fixed character counts alone.
  • Rerank before inserting, and cap the number of chunks.
  • Include the source path or URL with each chunk so the model can cite and you can debug.

Measure it

Context engineering is not guesswork if you evaluate it. Keep a small set of representative tasks and track success rate, tokens per task and steps per task as you change prompts, tools and retrieval. A change that cuts tokens by 40% with no drop in success is a real win.

Key takeaways

  • The context window is a budget, not a bucket.
  • Write, select, compress and isolate cover most techniques.
  • Tool definitions and tool output are the most common sources of bloat.
  • Measure success and token cost together, on a fixed task set.