Token Optimization and Context Engineering
Token optimization means reducing waste while preserving answer quality, safety, and verification. Tokens are model-specific units, so estimates below are illustrative rather than a universal tokenizer count.
Token fundamentals and cost accounting
| Measure | Meaning | Common accounting mistake |
|---|---|---|
| Input tokens | Instructions, user prompt, files, history, tool results, retrieved material | Counting only the user message |
| Output tokens | Model-generated visible or structured response | Ignoring verbose intermediate worker reports |
| Reasoning tokens | Model-specific internal/reasoning usage where surfaced by a provider | Assuming all providers expose or price it identically |
| Cached tokens | Eligible repeated input handled via provider cache, if supported | Treating all repeated text as free or cached |
| Context window | Active request capacity and output allowance | Filling the window so there is no room to respond or use tools |
| Budget | Cost/capacity reserved for this task or run | Tracking only a nominal maximum, not actual usage and retries |
Context overflow can truncate useful evidence, fail requests, or crowd out output. Cost estimation needs provider-specific prices and definitions; keep unit prices configurable and dated. Also measure latency, retries, model repair rate, and human review—not tokens alone.
estimated_cost = Σroute[(input_tokens × input_rate) + (output_tokens × output_rate)
+ cached_tokens × cached_rate + tool/compute charges]
orchestration_tokens ≈ root_context + Σworkers(worker_context + worker_output)
+ integration_context + validation/retry_contextThis is accounting notation, not a provider bill formula: rates, cache discounts, reasoning usage, and tool charges vary by product and can change.
def orchestration_tokens(root_input, worker_inputs, worker_outputs, integration, validation):
return root_input + sum(worker_inputs) + sum(worker_outputs) + integration + validation
example = orchestration_tokens(8000, [5000, 4000], [700, 500], 1800, 1200)The calculation deliberately adds measured or estimated usage from every phase; production accounting should also preserve per-model and cached-token dimensions.
Optimization techniques
| Technique | How / when it helps | Downside and accuracy guard | Likely cost / quality effect |
|---|---|---|---|
| Task-specific context | Send only constraints, relevant files/symbols, expected output | Missing a hidden dependency; include dependency map and uncertainty | Less input and irrelevant evidence; can lower correctness if scope is too narrow |
| Progressive disclosure | Start with index/symbol search, read a file/range only when relevant | Navigation mistakes; retain a way to expand on evidence | Smaller initial call, possibly more search calls; usually improves focus |
| Symbol/code search | Locate declarations/call sites before opening source | Text search can miss dynamic/reflection behavior; use AST/index where available | Low context cost; quality depends on recall and tool/index freshness |
| Incremental file reading | Read relevant ranges, then nearby call sites/tests | May miss whole-file invariants; inspect imports, callers, and tests | Avoids unrelated file tokens; preserves accuracy when dependency checks follow |
| Diff-based analysis | Review changed lines plus surrounding contracts | Existing code can be the source of a bug; inspect impacted dependencies as needed | Reduces review input; too little baseline context creates false negatives |
| Context pruning | Remove stale history and repeated tool output after extracting decisions/evidence | May discard a constraint needed later; retain a compact invariant list and source IDs | Less repeated input; quality stays stable only if critical facts survive |
| Selective repository indexing | Index stable source metadata/symbols, not every generated artifact | Stale indexes; invalidate on branch and file changes | Smaller retrieval payload; stale symbol maps can misdirect work |
| Context summarization | Compress investigation into facts, evidence, caveats, and open questions | Loses nuance; keep source paths/line references and confidence | Saves repeated context; quality depends on faithful, provenance-rich summaries |
| Retrieval-based context | Fetch relevant docs by query and ACL | Retrieval misses or injects stale/malicious text; version/filter sources | Replaces broad document input with focused passages; recall/grounding may improve or regress |
| Prompt compression | Remove repeated prose and irrelevant examples | Over-compressed instructions become ambiguous; test completion/error rate | Lowers input tokens; ambiguity may increase repair/retry cost |
| Structured responses | JSON/schema or fixed headings reduce integration overhead | Rigid schema can omit useful uncertainty; include risks/evidence fields | Often lowers integration/output waste; validation improves consistency, not truth |
| Output length control | Ask for concise result plus links to evidence | Could omit rationale; specify what must not be omitted | Reduces output tokens; short answers need traceable evidence for review |
| Provider prompt-prefix caching | Reuse eligible identical stable prefixes when supported | Eligibility, retention, privacy, and billing are provider-specific; measure cache hits | Can lower repeated-prefix processing cost/latency; does not improve accuracy by itself |
| Semantic caching | Reuse prior response for equivalent request classes | Similarity is not equivalence; gate on tenant, freshness, policy, and confidence | May avoid full model calls; stale/cross-tenant answers can make quality and risk unacceptable |
| Model routing | Match model capability/cost to task | Misrouting or a weaker model can increase repair cost; evaluate end-to-end | Can lower average cost; route errors can lower quality or raise retries |
| Avoid repeated tool calls | Cache deterministic read results within one snapshot | Results can go stale; bind cache to commit/version and invalidate on edits | Saves tool latency/cost; stale results damage correctness |
| Concurrency limits | Avoid many agents repeating broad context | Less parallel speed; reserve concurrency for independent work | Reduces duplicated calls/context; may increase wall-clock latency |
For every optimization, evaluate both tokens and task quality. Use a held-out benchmark and compare successful completion, regressions, false confidence, latency, and cost per accepted result.
Context engineering and trust boundaries
Assemble context from explicit layers, in order of authority and relevance:
- System/developer policy and tool permission envelope.
- User goal and constraints.
- Project instructions relevant to the touched paths.
- Repository facts: commit, symbol map, targeted source/tests, relevant diff.
- Retrieved documentation with source, version/date, and access labels.
- Tool results and prior agent reports, clearly marked as untrusted observations.
- Task memory: decisions, evidence, unresolved questions, and next step.
Never merge a README, webpage, tool output, or subagent response into the instruction layer. Repository content can contain accidental or malicious instructions; treat it as data to analyze. Keep task context isolated by project/tenant, redact secrets before model access, remove stale facts, and preserve provenance. Summaries should say “observed in file/version” rather than turn claims into authoritative instructions.
Context package example
{
"goal": "Add loading cancellation to the profile screen",
"commit": "<current-commit-id>",
"owned_paths": ["App/Profile/ProfileViewModel.swift", "Tests/ProfileTests.swift"],
"read_only_references": ["App/API/ProfileClient.swift", "App/Profile/ProfileView.swift"],
"constraints": ["Preserve existing API contract", "Do not log profile payloads"],
"acceptance": ["Cancel stale request", "Test completion and cancellation"],
"untrusted_observations": ["Current implementation uses task cancellation in the view"],
"excluded": ["DerivedData", "build output", "unrelated feature modules"]
}This compact structure provides task boundaries, not a substitute for repository-level authorization or validating the commit/path state.
Repository-aware coding agents
Reading every file is wasteful when most files are unrelated and can introduce more stale or adversarial text. Begin with the current directory map, build/test instructions, git status/diff, and project instructions. Search symbols and references; then inspect relevant declarations, callers, tests, protocol implementations, generated-code boundaries, and configuration that changes behavior. Use AST or language-server indexing for accurate symbol references where available; text search remains useful but is not complete for dynamic dispatch, macros, generated code, or reflection.
For Swift, search the view/model/repository protocol and test double, then inspect target membership, actor isolation, availability, and call sites. For TypeScript, trace exported symbol → route/schema → service → repository/query and its tests; inspect package scripts and lockfile rather than every dependency. Before editing, record in-scope paths. After editing, inspect git diff, run the smallest relevant test first, then broader validation based on risk. Do not assume code search results are complete without checking module boundaries and generated source.
Illustrative repository context budget
Assumptions for comparison only: a hypothetical repository has 200 files averaging 1,500 tokens of source each (=300,000 source tokens before tool metadata); a targeted five-file package averages 1,500 tokens/file (=7,500); model/tool overhead, history, summaries, and output are excluded. Actual tokenizer and source sizes vary.
| Scenario | Approximate source loaded | Why it changes |
|---|---|---|
| A. One agent reads all 200 files | ~300,000 tokens | Broad reading before one screen change, likely context overflow or repeated compaction |
| B. Search, then read five files | ~7,500 tokens | Focused call graph plus tests; expand if evidence shows hidden dependencies |
| C. Five workers each read all 200 | ~1,500,000 aggregate tokens | Parallel duplication; can also increase provider/cache costs and integration output |
| D. Supervisor sends a 5-file task package to each of 5 different owners | up to ~37,500 aggregate source tokens if packages are disjoint | Tasks share only necessary contracts; add supervisor/integration context and worker outputs |
These are not universal savings claims and do not compare reasoning or answer quality. Scenario D is sensible only if the tasks are truly independent and file ownership is non-overlapping. If the workers all need the same source, measure whether shared context/cache or one sequential agent is cheaper.
Subagent token efficiency
| Strategy | Duplicated context | Coordination cost | Typical fit |
|---|---|---|---|
| Direct execution | None across agents | Low | Small or dependent work |
| One subagent | One task package + handoff | Moderate | Independent review/research |
| Sequential subagents | Repeated summaries/next-task context | High latency and handoff | Strict phase gates with specialist handoff |
| Parallel subagents | Shared repo context can repeat; outputs merge | High integration | Independent bounded tasks |
| Hierarchical | Context can duplicate at each level | Highest | Large work with genuine domain decomposition |
Budget formula: total = root input/output + Σ(worker input/output + tool results) + integration + validation + retries. Record actual token usage by run, task, route, and cache status where available. Estimate total cost with the provider's current pricing, not a hard-coded guide number.
Token budget manager
Reserve before delegation, account actual usage after each call, and never allow workers to borrow an unlimited shared pool. Protect a validation/report reserve so a task does not exhaust its budget before checking output. Prioritize correctness/security gates over optional polish; if budget is exhausted, stop or request a model/budget decision rather than silently weakening quality.
type Budget = { input: number; output: number; calls: number };
type Usage = Budget;
class BudgetManager {
private reserved = new Map<string, Budget>();
private spent: Budget = { input: 0, output: 0, calls: 0 };
constructor(private readonly limit: Budget, private readonly validationReserve: Budget) {}
reserve(taskId: string, amount: Budget): void {
const current = [...this.reserved.values()].reduce(add, this.spent);
const ceiling = subtract(this.limit, this.validationReserve);
if (!fits(add(current, amount), ceiling)) throw new Error("Budget reservation denied");
this.reserved.set(taskId, amount);
}
settle(taskId: string, actual: Usage): void {
const held = this.reserved.get(taskId);
if (!held) throw new Error("No reservation for task");
this.reserved.delete(taskId);
this.spent = add(this.spent, actual); // production code validates non-negative values
if (!fits(this.spent, this.limit)) throw new Error("Actual usage exceeded global budget");
}
}
function add(a: Budget, b: Budget): Budget {
return { input: a.input + b.input, output: a.output + b.output, calls: a.calls + b.calls };
}
function subtract(a: Budget, b: Budget): Budget {
return { input: a.input - b.input, output: a.output - b.output, calls: a.calls - b.calls };
}
function fits(a: Budget, limit: Budget): boolean {
return a.input <= limit.input && a.output <= limit.output && a.calls <= limit.calls;
}This is illustrative single-process pseudocode. A production manager needs atomic durable reservations, model-specific input/output rates, provider usage reconciliation, concurrency safety, currency limits, partial usage on failure, and escalation rules. Token limits do not equal dollar limits.
{
"task_budget": { "input_tokens": 24000, "output_tokens": 6000, "tool_calls": 18 },
"validation_reserve": { "input_tokens": 5000, "output_tokens": 1500, "tool_calls": 4 },
"per_agent": {
"architecture_review": { "input_tokens": 5000, "output_tokens": 1200, "tool_calls": 2 },
"implementation": { "input_tokens": 10000, "output_tokens": 2500, "tool_calls": 10 },
"security_review": { "input_tokens": 4000, "output_tokens": 1000, "tool_calls": 2 }
},
"on_exhaustion": "stop_and_report_with_evidence",
"human_approval_required_for": ["external_write", "deployment", "destructive_change"]
}Numbers are a sample configuration, not recommended universal budgets.
Model routing by quality-per-successful-task
Route a short deterministic extraction to a small model only after evaluation. Route difficult reasoning, high-impact design, or tool-plan ambiguity to a stronger approved model when its latency and privacy characteristics fit. Use local models for data/control constraints only after measuring task quality and operational cost. If confidence is low, abstain, retrieve more, or ask a human; automatic escalation must respect the same budget and data policy.
| Route class | Candidate | Gate before accepting |
|---|---|---|
| Classification / normalized extraction | Small fast model | Schema parse + labeled accuracy threshold |
| Summarization of trusted material | Small or general model | Source coverage, omission review, privacy filter |
| Coding change with local tests | Coding-capable model | Tests, diff constraints, code review |
| Architecture/security review | Strong reasoning model + independent deterministic checks | Evidence, threat checklist, human review for high risk |
| Sensitive/high-compliance task | Approved endpoint/local deployment | Data residency, access, retention, and model evaluation approval |
Context/token interview hit points
🧠 Say: “I minimize irrelevant context by locating symbols and dependencies, but I keep the contract, tests, instructions, and enough surrounding code to verify behavior. I judge token savings by cost per accepted, correct change—not by a smaller prompt.”
