Token Optimization

Manage model context, token budgets, retrieval, caching, routing, and cost-quality tradeoffs.

~20 min reading
Suggest content

Token Optimization and Context Engineering

Token optimization means reducing waste while preserving answer quality, safety, and verification. Tokens are model-specific units, so estimates below are illustrative rather than a universal tokenizer count.

Token fundamentals and cost accounting

MeasureMeaningCommon accounting mistake
Input tokensInstructions, user prompt, files, history, tool results, retrieved materialCounting only the user message
Output tokensModel-generated visible or structured responseIgnoring verbose intermediate worker reports
Reasoning tokensModel-specific internal/reasoning usage where surfaced by a providerAssuming all providers expose or price it identically
Cached tokensEligible repeated input handled via provider cache, if supportedTreating all repeated text as free or cached
Context windowActive request capacity and output allowanceFilling the window so there is no room to respond or use tools
BudgetCost/capacity reserved for this task or runTracking only a nominal maximum, not actual usage and retries

Context overflow can truncate useful evidence, fail requests, or crowd out output. Cost estimation needs provider-specific prices and definitions; keep unit prices configurable and dated. Also measure latency, retries, model repair rate, and human review—not tokens alone.

TEXT
estimated_cost = Σroute[(input_tokens × input_rate) + (output_tokens × output_rate)
                        + cached_tokens × cached_rate + tool/compute charges]
​
orchestration_tokens ≈ root_context + Σworkers(worker_context + worker_output)
                     + integration_context + validation/retry_context

This is accounting notation, not a provider bill formula: rates, cache discounts, reasoning usage, and tool charges vary by product and can change.

PYTHON
def orchestration_tokens(root_input, worker_inputs, worker_outputs, integration, validation):
    return root_input + sum(worker_inputs) + sum(worker_outputs) + integration + validation
​
example = orchestration_tokens(8000, [5000, 4000], [700, 500], 1800, 1200)

The calculation deliberately adds measured or estimated usage from every phase; production accounting should also preserve per-model and cached-token dimensions.

Optimization techniques

TechniqueHow / when it helpsDownside and accuracy guardLikely cost / quality effect
Task-specific contextSend only constraints, relevant files/symbols, expected outputMissing a hidden dependency; include dependency map and uncertaintyLess input and irrelevant evidence; can lower correctness if scope is too narrow
Progressive disclosureStart with index/symbol search, read a file/range only when relevantNavigation mistakes; retain a way to expand on evidenceSmaller initial call, possibly more search calls; usually improves focus
Symbol/code searchLocate declarations/call sites before opening sourceText search can miss dynamic/reflection behavior; use AST/index where availableLow context cost; quality depends on recall and tool/index freshness
Incremental file readingRead relevant ranges, then nearby call sites/testsMay miss whole-file invariants; inspect imports, callers, and testsAvoids unrelated file tokens; preserves accuracy when dependency checks follow
Diff-based analysisReview changed lines plus surrounding contractsExisting code can be the source of a bug; inspect impacted dependencies as neededReduces review input; too little baseline context creates false negatives
Context pruningRemove stale history and repeated tool output after extracting decisions/evidenceMay discard a constraint needed later; retain a compact invariant list and source IDsLess repeated input; quality stays stable only if critical facts survive
Selective repository indexingIndex stable source metadata/symbols, not every generated artifactStale indexes; invalidate on branch and file changesSmaller retrieval payload; stale symbol maps can misdirect work
Context summarizationCompress investigation into facts, evidence, caveats, and open questionsLoses nuance; keep source paths/line references and confidenceSaves repeated context; quality depends on faithful, provenance-rich summaries
Retrieval-based contextFetch relevant docs by query and ACLRetrieval misses or injects stale/malicious text; version/filter sourcesReplaces broad document input with focused passages; recall/grounding may improve or regress
Prompt compressionRemove repeated prose and irrelevant examplesOver-compressed instructions become ambiguous; test completion/error rateLowers input tokens; ambiguity may increase repair/retry cost
Structured responsesJSON/schema or fixed headings reduce integration overheadRigid schema can omit useful uncertainty; include risks/evidence fieldsOften lowers integration/output waste; validation improves consistency, not truth
Output length controlAsk for concise result plus links to evidenceCould omit rationale; specify what must not be omittedReduces output tokens; short answers need traceable evidence for review
Provider prompt-prefix cachingReuse eligible identical stable prefixes when supportedEligibility, retention, privacy, and billing are provider-specific; measure cache hitsCan lower repeated-prefix processing cost/latency; does not improve accuracy by itself
Semantic cachingReuse prior response for equivalent request classesSimilarity is not equivalence; gate on tenant, freshness, policy, and confidenceMay avoid full model calls; stale/cross-tenant answers can make quality and risk unacceptable
Model routingMatch model capability/cost to taskMisrouting or a weaker model can increase repair cost; evaluate end-to-endCan lower average cost; route errors can lower quality or raise retries
Avoid repeated tool callsCache deterministic read results within one snapshotResults can go stale; bind cache to commit/version and invalidate on editsSaves tool latency/cost; stale results damage correctness
Concurrency limitsAvoid many agents repeating broad contextLess parallel speed; reserve concurrency for independent workReduces duplicated calls/context; may increase wall-clock latency

For every optimization, evaluate both tokens and task quality. Use a held-out benchmark and compare successful completion, regressions, false confidence, latency, and cost per accepted result.

Context engineering and trust boundaries

Assemble context from explicit layers, in order of authority and relevance:

  1. System/developer policy and tool permission envelope.
  2. User goal and constraints.
  3. Project instructions relevant to the touched paths.
  4. Repository facts: commit, symbol map, targeted source/tests, relevant diff.
  5. Retrieved documentation with source, version/date, and access labels.
  6. Tool results and prior agent reports, clearly marked as untrusted observations.
  7. Task memory: decisions, evidence, unresolved questions, and next step.

Never merge a README, webpage, tool output, or subagent response into the instruction layer. Repository content can contain accidental or malicious instructions; treat it as data to analyze. Keep task context isolated by project/tenant, redact secrets before model access, remove stale facts, and preserve provenance. Summaries should say “observed in file/version” rather than turn claims into authoritative instructions.

Context package example

JSON
{
  "goal": "Add loading cancellation to the profile screen",
  "commit": "<current-commit-id>",
  "owned_paths": ["App/Profile/ProfileViewModel.swift", "Tests/ProfileTests.swift"],
  "read_only_references": ["App/API/ProfileClient.swift", "App/Profile/ProfileView.swift"],
  "constraints": ["Preserve existing API contract", "Do not log profile payloads"],
  "acceptance": ["Cancel stale request", "Test completion and cancellation"],
  "untrusted_observations": ["Current implementation uses task cancellation in the view"],
  "excluded": ["DerivedData", "build output", "unrelated feature modules"]
}

This compact structure provides task boundaries, not a substitute for repository-level authorization or validating the commit/path state.

Repository-aware coding agents

Reading every file is wasteful when most files are unrelated and can introduce more stale or adversarial text. Begin with the current directory map, build/test instructions, git status/diff, and project instructions. Search symbols and references; then inspect relevant declarations, callers, tests, protocol implementations, generated-code boundaries, and configuration that changes behavior. Use AST or language-server indexing for accurate symbol references where available; text search remains useful but is not complete for dynamic dispatch, macros, generated code, or reflection.

For Swift, search the view/model/repository protocol and test double, then inspect target membership, actor isolation, availability, and call sites. For TypeScript, trace exported symbol → route/schema → service → repository/query and its tests; inspect package scripts and lockfile rather than every dependency. Before editing, record in-scope paths. After editing, inspect git diff, run the smallest relevant test first, then broader validation based on risk. Do not assume code search results are complete without checking module boundaries and generated source.

Illustrative repository context budget

Assumptions for comparison only: a hypothetical repository has 200 files averaging 1,500 tokens of source each (=300,000 source tokens before tool metadata); a targeted five-file package averages 1,500 tokens/file (=7,500); model/tool overhead, history, summaries, and output are excluded. Actual tokenizer and source sizes vary.

ScenarioApproximate source loadedWhy it changes
A. One agent reads all 200 files~300,000 tokensBroad reading before one screen change, likely context overflow or repeated compaction
B. Search, then read five files~7,500 tokensFocused call graph plus tests; expand if evidence shows hidden dependencies
C. Five workers each read all 200~1,500,000 aggregate tokensParallel duplication; can also increase provider/cache costs and integration output
D. Supervisor sends a 5-file task package to each of 5 different ownersup to ~37,500 aggregate source tokens if packages are disjointTasks share only necessary contracts; add supervisor/integration context and worker outputs

These are not universal savings claims and do not compare reasoning or answer quality. Scenario D is sensible only if the tasks are truly independent and file ownership is non-overlapping. If the workers all need the same source, measure whether shared context/cache or one sequential agent is cheaper.

Subagent token efficiency

StrategyDuplicated contextCoordination costTypical fit
Direct executionNone across agentsLowSmall or dependent work
One subagentOne task package + handoffModerateIndependent review/research
Sequential subagentsRepeated summaries/next-task contextHigh latency and handoffStrict phase gates with specialist handoff
Parallel subagentsShared repo context can repeat; outputs mergeHigh integrationIndependent bounded tasks
HierarchicalContext can duplicate at each levelHighestLarge work with genuine domain decomposition

Budget formula: total = root input/output + Σ(worker input/output + tool results) + integration + validation + retries. Record actual token usage by run, task, route, and cache status where available. Estimate total cost with the provider's current pricing, not a hard-coded guide number.

Token budget manager

Reserve before delegation, account actual usage after each call, and never allow workers to borrow an unlimited shared pool. Protect a validation/report reserve so a task does not exhaust its budget before checking output. Prioritize correctness/security gates over optional polish; if budget is exhausted, stop or request a model/budget decision rather than silently weakening quality.

TYPESCRIPT
type Budget = { input: number; output: number; calls: number };
type Usage = Budget;
​
class BudgetManager {
  private reserved = new Map<string, Budget>();
  private spent: Budget = { input: 0, output: 0, calls: 0 };
​
  constructor(private readonly limit: Budget, private readonly validationReserve: Budget) {}
​
  reserve(taskId: string, amount: Budget): void {
    const current = [...this.reserved.values()].reduce(add, this.spent);
    const ceiling = subtract(this.limit, this.validationReserve);
    if (!fits(add(current, amount), ceiling)) throw new Error("Budget reservation denied");
    this.reserved.set(taskId, amount);
  }
​
  settle(taskId: string, actual: Usage): void {
    const held = this.reserved.get(taskId);
    if (!held) throw new Error("No reservation for task");
    this.reserved.delete(taskId);
    this.spent = add(this.spent, actual); // production code validates non-negative values
    if (!fits(this.spent, this.limit)) throw new Error("Actual usage exceeded global budget");
  }
}
​
function add(a: Budget, b: Budget): Budget {
  return { input: a.input + b.input, output: a.output + b.output, calls: a.calls + b.calls };
}
function subtract(a: Budget, b: Budget): Budget {
  return { input: a.input - b.input, output: a.output - b.output, calls: a.calls - b.calls };
}
function fits(a: Budget, limit: Budget): boolean {
  return a.input <= limit.input && a.output <= limit.output && a.calls <= limit.calls;
}

This is illustrative single-process pseudocode. A production manager needs atomic durable reservations, model-specific input/output rates, provider usage reconciliation, concurrency safety, currency limits, partial usage on failure, and escalation rules. Token limits do not equal dollar limits.

JSON
{
  "task_budget": { "input_tokens": 24000, "output_tokens": 6000, "tool_calls": 18 },
  "validation_reserve": { "input_tokens": 5000, "output_tokens": 1500, "tool_calls": 4 },
  "per_agent": {
    "architecture_review": { "input_tokens": 5000, "output_tokens": 1200, "tool_calls": 2 },
    "implementation": { "input_tokens": 10000, "output_tokens": 2500, "tool_calls": 10 },
    "security_review": { "input_tokens": 4000, "output_tokens": 1000, "tool_calls": 2 }
  },
  "on_exhaustion": "stop_and_report_with_evidence",
  "human_approval_required_for": ["external_write", "deployment", "destructive_change"]
}

Numbers are a sample configuration, not recommended universal budgets.

Model routing by quality-per-successful-task

Route a short deterministic extraction to a small model only after evaluation. Route difficult reasoning, high-impact design, or tool-plan ambiguity to a stronger approved model when its latency and privacy characteristics fit. Use local models for data/control constraints only after measuring task quality and operational cost. If confidence is low, abstain, retrieve more, or ask a human; automatic escalation must respect the same budget and data policy.

Route classCandidateGate before accepting
Classification / normalized extractionSmall fast modelSchema parse + labeled accuracy threshold
Summarization of trusted materialSmall or general modelSource coverage, omission review, privacy filter
Coding change with local testsCoding-capable modelTests, diff constraints, code review
Architecture/security reviewStrong reasoning model + independent deterministic checksEvidence, threat checklist, human review for high risk
Sensitive/high-compliance taskApproved endpoint/local deploymentData residency, access, retention, and model evaluation approval

Context/token interview hit points

🧠 Say: “I minimize irrelevant context by locating symbols and dependencies, but I keep the contract, tests, instructions, and enough surrounding code to verify behavior. I judge token savings by cost per accepted, correct change—not by a smaller prompt.”

References