Context Engineering
Context is a working set assembled for each step
Model input is not a direct copy of database records. At the start of a run, aibuddy builds system instructions and the tool set, then reconstructs the persisted reduction view over raw task messages. Inside the agent loop, each step decides whether to retain that view, reveal another capability, or reduce older material.
This boundary separates what the system owns from what the model needs now. Task records retain facts, the runtime produces the current view, and the model receives the working set required for its next action.
Stable prefixes and dynamic tails have different jobs
System instructions are ordered by change rate. Stable content comes first and volatile content later, preserving a reusable prompt-cache prefix without hiding current environment state.
| Order | Content | Lifetime |
|---|---|---|
| 1 | Base agent instructions and agent-specific rules | Usually stable within a task |
| 2 | Built-in tool instructions and scenario workflow | Fixed by the run configuration |
| 3 | Memory index | Stable within a turn and present only with the memory tool |
| 4 | Date, timezone, runtime, and workspace | May change at each assembly |
| Message tail | Plan, todos, team state, and one-shot reminders | Read immediately before each model call |
Instructions and built-in tools are constructed concurrently, then merged with MCP tools. Date and workspace form the final system section, while in-turn reminders enter as a separate tail input. Updating either does not require rewriting the base rules.
Caching guides ordering but does not override correctness. Identity, permission, scenario, and environment changes must reach the model even when they invalidate a cache prefix.
Progressive disclosure occurs at capability boundaries
aibuddy applies one principle to Skills, MCP, memory, and knowledge: expose enough information to judge relevance before loading the body required for execution. Their loading paths remain distinct.
| Capability | Initially visible | Deeper load |
|---|---|---|
| Skill | Activation name and load directive | view_skill reads SKILL.md, then referenced material |
| MCP | Small tool sets are direct; large sets retain a search surface | Native tool search or the activeTools path activates matching schemas |
| Memory | Titles and compact index | memory_recall selects and retrieves relevant memories |
| Knowledge | The main agent receives a delegable knowledge capability | A scoped knowledge subagent reads the document map, outline, and bounded sections |
| Large tool result | Reduced record and recovery hint | view_tool_call or a tool-native path retrieves source output |
MCP deferral is independent of the model window. It starts when one server exposes more than 10 tools or the aggregate exceeds 30. Native Anthropic and OpenAI search paths keep the tool array stable; other providers use tool_search with activeTools as the portable fallback.
Knowledge retrieval adds another isolation boundary. The knowledge subagent is offered only when usable documents are attached to the task. It begins with the document map or search_docs, opens structure through read_doc, and pages bounded content through read_section. The main context receives a synthesized answer and citations instead of the intermediate navigation trace.
Progressive disclosure reduces resident tokens and selection noise, but it can add a discovery step. Discovery failures and execution failures require different diagnosis: the former points to indexes, queries, or scope; the latter to parameters, environment, or implementation.
In-turn state does not require rereading the transcript
Plans, todos, team-member state, and temporary constraints should not be reconstructed from a long transcript at every step. RuntimeContext stores reminders in keyed slots. Each producer updates its own value without overwriting other producers or accumulating stale versions.
The runtime reads the current slots before a model call and injects them at the message tail. One-shot keys prefixed with once: are removed after delivery; baseline and still-active state remain available. These slots are turn-scoped scratch state, not a replacement for persisted task messages, status, or deliverables.
Tail placement also means a preference or progress update does not mutate the stable system prefix. Reminders remain model input, however; authorization, task locking, and completion conditions are independently enforced by the runtime.
Compaction creates a view without rewriting task records
The budget first subtracts instructions, tool schemas, and reserved output from the model window. The default policy starts reduction at 65% of the effective window and targets 50%, leaving hysteresis so adjacent steps do not compact repeatedly.
Reduction proceeds in this order:
- A new turn reconstructs raw messages and replays persisted snapshots plus the previous turn’s measured usage anchor.
- Tool results larger than 2,500 tokens are precompacted in the background when execution ends.
- At the trigger, the runtime reduces older large tool results first while protecting the latest five execution steps.
- If the prompt remains over budget, it summarizes an older message prefix into task context.
- Two consecutive attempts saving less than 10% latch further retries until the input grows materially.
Snapshots are stored separately under stable addresses, so raw task messages are not overwritten. Recoverable tool output retains a view_tool_call or tool-native retrieval hint; only explicitly disposable intermediate output collapses to a minimal terminal record. Compaction is therefore a reconstructable projection, not a destructive edit of the source record.
Isolation can be safer than further compaction
Knowledge retrieval naturally involves repeated search, outline inspection, and section reads, so aibuddy places that trace in a knowledge subagent. General subagents also run bounded research or execution in independent contexts and return conclusions, status, and deliverables to the parent.
Isolation is not free compression. Delegated objectives must be self-contained, and returned results can omit intermediate evidence. It fits branches with explicit acceptance criteria; work that depends tightly on the parent trajectory should stay in the current context.
Diagnose the assembly stage from the symptom
| Symptom | Inspect first |
|---|---|
| Attachments, memory, or knowledge never affect the answer | Authorization scope, enabled capability, index, and discovery result |
| The agent selects the wrong tool | Resident tool count, MCP deferral threshold, and tool descriptions |
| A plan becomes a claim of completion | Dynamic-state producer and compaction summary |
| Source values or identifiers disappear | Tool compactor, recovery hint, and summary coverage |
| Cost grows on every step | Stable-prefix churn, compaction threshold, and tool-schema overhead |
| Compaction repeats with little benefit | Ineffective-attempt count and volume of new content |
Follow selection, assembly, budget, then recovery. Adding more system instructions usually increases the input without correcting missing authorization, a poor index, or an unrecoverable tool result.
Implementation anchors
| Mechanism | Primary implementation |
|---|---|
| System instruction order | instructionsBuilder.build in instructions-builder.ts |
| Tool and instruction assembly | buildAgentInputs in agent-loop.ts |
| Keyed dynamic reminders | RuntimeContext.reminders in context.ts |
| Skill activation | inject-user-skills.ts and view_skill |
| MCP deferral | agent-setup.ts, tool-defer.ts, and tool-discovery.ts |
| Compaction budget and ladder | budget.ts, compact.ts, and compaction-state.ts |
| Tool-result recovery | tool-compactor.ts and view_tool_call |
| Progressive knowledge reads | knowledge-subagent.ts, read_doc, search_docs, and read_section |
Related reading
- Prompt System for stable instruction structure and cache boundaries.
- Agent Status Bar for directly available plan and todo state.
- Context Compaction for reduction, snapshots, and recovery.
- Memory & Knowledge for cross-task memory and knowledge lifetimes.