AI Engineering

Design LLM applications, retrieval, model routing, agent orchestration, evaluation, and production AI workflows.

~21 min reading
Suggest content

AI Engineering and Orchestration

Use this sheet for AI application architecture and orchestration concepts. Tool-specific behavior belongs in AI Coding Tools and MCP.

AI and LLM fundamentals

Artificial intelligence is the broad field of systems performing tasks associated with perception, reasoning, planning, or generation. Machine learning fits behavior from data; a neural network is a parameterized function built from connected computational layers; deep learning uses many such layers. A transformer is a neural architecture built around attention. A large language model (LLM) predicts/generates token sequences from learned parameters and the supplied context. It does not automatically know whether a response is true, current, authorized, or grounded in a source.

TermPractical definitionEngineering implication
TokenizationConverts text/code into model-specific unitsCharacter counts do not reliably estimate token counts
Context windowMaximum active input plus generated token space for a model/requestBudget room for instructions, tools, history, and output
AttentionMechanism for relating representations in the provided sequenceMore context is not automatically better context
EmbeddingVector representation useful for semantic similaritySimilarity is a retrieval signal, not proof or authorization
InferenceRunning a trained model to produce outputsLatency/cost depend on model, input/output size, and serving
Training / fine-tuningUpdating model parameters on data or examplesIt is different from providing documents at inference time
QuantizationLower-precision representation to reduce model size/computeCan trade quality for memory/latency; measure the target task
HallucinationPlausible output that is unsupported or incorrectValidate important claims and constrain side effects
EvaluationRepeatable measurement against task-specific casesUse quality, safety, cost, latency, and regression metrics

Prompt engineering improves instruction clarity. It cannot guarantee truth or security. Prefer constrained schemas, trusted retrieval, deterministic checks, and human approval for consequential actions.

LLM application architecture

TEXT
iOS client → authenticated API → policy / quota / request validation
                              → prompt + context builder
                              → model router / LLM provider
                               ↘ tools (scoped, validated, audited)
                                ↘ retrieval (ACL filter → search → rerank)
                              → output validation → stream / response
                              → traces, evaluation, usage and cost records

Function/tool calling is a model-proposed structured action, not execution authorization. Validate schema and business policy, enforce the calling user's permissions at the tool boundary, execute with least privilege, and return a minimal result. Structured output constrains shape, not truth. Streaming improves perceived latency but partial text is not a completed/validated answer. Conversation context should be deliberately selected, compacted, and access-filtered.

Retrieval-Augmented Generation (RAG)

TEXT
Authorized source → parse → chunk + metadata → embed/index
User query → authenticate → ACL filter → keyword + vector search
          → rerank → bounded passages with source IDs → model answer
          → citation/entailment checks → response + trace

Use RAG when answers depend on changing, private, or attributable external material. Chunk around semantic boundaries; retain document/version/tenant/access metadata; filter by authorization before exposing content; tune lexical and vector retrieval; rerank only if evaluation shows benefit. Cite retrieved sources and allow “not enough evidence.” A vector store must not be the authorization system. Evaluate retrieval recall, answer groundedness, citation correctness, latency, and permission leakage with adversarial and stale-document cases.

Prompt caching, semantic caching, and response freshness

Prompt caching may reuse processing for a repeated eligible prompt prefix when the provider/model supports it. Put stable instructions/schema before dynamic user data where that matches the provider's cache rules; measure actual cache hits and latency/cost. Eligibility, minimum prefix sizes, retention, and billing are provider-specific and change over time.

Semantic caching stores a prior answer for a query judged similar to a new one. It is not prompt caching and carries application-level correctness risks: bind keys to user/tenant/authorization scope, model/prompt version, retrieval corpus version, locale, and freshness policy. Do not reuse personalized or security-sensitive answers based on vector similarity alone. Invalidate on source, permission, policy, or model changes; use an exact deterministic cache only when semantics permit it.

Model selection and routing

Model selection is a workload decision, not a permanent leaderboard. Compare candidate models on representative prompts and blinded evaluation data.

SignalSmall / fast model may fitStronger reasoning / specialized model may fit
ComplexityClassification, extraction, short transformationMulti-constraint planning, subtle debugging, architecture review
LatencyInline autocomplete or high-volume triageUser accepts a longer analysis step
Tool useStable, narrow schema with bounded actionAmbiguous planning or more complex tool selection
PrivacyLocally hosted or approved endpoint is requiredOnly if deployment/data policy permits it
ContextRelevant input fits with headroomLong documents/code require larger validated capacity
CostRepetitive low-risk requestsCost justified by measured success/repair rate
Model family / deploymentLocal model, small general model, or cloud model can satisfy policy and benchmarkCoding-specialized model or stronger reasoning model wins on the measured task and is permitted for its data

Route on task type, risk, context size, tool needs, quality threshold, latency SLA, and privacy policy. The cheapest call may be expensive overall if it requires retries, human repair, or causes an incorrect action. Use a fallback with explicit downgrade behavior, preserve evaluation labels, and monitor per-route quality/cost. Avoid hard-coded model rankings; model names and limits change.

Agent, tool, skill, workflow, and orchestrator

ConceptWhat it isExampleKey distinction
ToolNarrow function/API with defined input and outputRead a file, search symbols, run testsExecutes one capability; not independently responsible for a goal
SkillReusable domain procedure/instructions, sometimes bundled with references or scriptsiOS code-review checklistGuides work; may use tools but is not necessarily an agent
WorkflowExplicit sequence/graph of steps and branchesValidate request → retrieve → summarize → approveControl flow is primarily predefined
AgentModel-driven loop that uses context/tools toward a bounded goalDiagnose failing test with read-only repo toolsChooses next step within its permissions; may be nondeterministic
OrchestratorCoordinates runs, context, budgets, permissions, state, and resultsSupervisor with bounded worker poolOwns system-level execution and final integration

Prefer a deterministic workflow when the steps and branch rules are known. Use an agent loop when decisions genuinely depend on observations. A longer loop is not inherently smarter or safer.

Orchestration patterns

Each pattern below is a design choice. Token costs and latency depend on context reuse, tool/model speed, concurrency, retries, and integration. Parallel calls reduce wall time only when tasks are independent and the critical path is not dominated by the slowest worker.

Single agent

TEXT
Request → agent ↔ scoped tools → validated result

One context owns planning and execution. Advantages: lowest coordination overhead, one coherent state, easy audit. Disadvantages: context/budget can grow; one model must cover every specialty. Cost/latency: usually least duplicated context; serial tool loop can be slow. Security/failure: one permission set limits blast radius, but injection or a bad plan can affect all exposed tools. Use: bounded task with a small tool surface.

Supervisor-worker

TEXT
                 ┌→ iOS worker ─┐
Request → supervisor → API worker ├→ validate → integrate
                 └→ test worker ─┘

Supervisor makes plan, assigns typed tasks, validates and merges. Advantages: domain focus and independent investigation. Disadvantages: delegation/integration and duplicated context. Cost/latency: more calls and tokens; parallel workers may lower wall time. Security/failure: workers receive narrow permissions; peer output remains untrusted data; supervisor can fail to detect bad work. Use: separable research/review or file-disjoint implementation.

Hierarchical multi-agent

TEXT
Root supervisor → domain lead(s) → specialist worker(s) → evidence

Hierarchy adds decomposers between root and experts. It can handle broad portfolios, but multiplies handoffs, context duplication, review distance, and failure propagation. Budget by total tree, not just leaves; use only when coordination complexity justifies the extra level. Limit depth, total calls, and delegated permissions.

Sequential execution

TEXT
Plan → inspect → implement → test → review → report

Each step consumes the previous result. Advantages: clear dependencies and low merge conflict. Disadvantages: serial latency and context accumulation. Security: validate the output before promoting it into trusted state. Failure: stop/repair at the failed gate rather than proceeding on a false assumption. Use: dependent phases or one shared file owner.

Parallel execution

TEXT
                    ┌→ independent API contract review ─┐
Plan → fan-out ─────┼→ threat review ────────────────────┼→ reconcile
                    └→ isolated test audit ──────────────┘

Good for independent analysis with explicit deliverables. It can increase cost and creates disagreement/merge work; wall time is bounded by slowest worker plus integration. Assign separate files or read-only tasks, cap concurrency, and cancel redundant work. Do not parallelize shared mutable state or unresolved decisions.

Router-based orchestration

TEXT
Request → policy router → [Swift | backend | security | general workflow]

Routes based on explicit signals or classifier. It reduces irrelevant context but may misroute or leak data to the wrong model. Keep a safe default, log route rationale, restrict sensitive inputs before routing, and measure route accuracy/latency/cost. Use when task classes map to known specialists.

Planner-executor

TEXT
Goal → bounded plan (steps, dependencies, exit criteria) → executor → check

Separates decomposition from action. Planning can improve complex work but creates extra model cost and stale plans when observations change. Revalidate assumptions after each step, enforce allowed action schemas, and stop at budget/approval boundaries. Use for multi-step work with observable completion criteria.

Evaluator-optimizer

TEXT
Candidate → deterministic/model evaluator → bounded revise loop → accept or stop

Improves drafts against a rubric. It can optimize a flawed metric, loop indefinitely, or reward style over truth. Use objective checks where possible, independent eval cases, maximum iterations, and human review for high-impact outputs. Track total retries in cost/latency.

Human-in-the-loop

TEXT
Agent prepares proposed change → policy gate → person reviews → scoped action

Use for irreversible, sensitive, privileged, externally visible, or low-confidence actions. The approval must show concrete diff/target/impact and expire if state changes. Human review adds latency; a generic “approve?” dialog with no context is not a meaningful control. Never rely on approval as a substitute for least privilege.

Event-driven agents

TEXT
Event → durable queue → dedupe/claim → bounded agent run → result/event

Supports asynchronous, long-running tasks and independent scaling. Requires idempotency, replay policy, dead-letter handling, quotas, event schema versioning, tracing, and cancellation semantics. Duplicate delivery is normal; do not promise exactly-once end-to-end behavior without explaining the transaction boundary. Use for background tasks, not as a default for simple interactive requests.

Engineering subagents and task contracts

Use a specialist only when the output can be independently specified and the gain exceeds delegation plus integration cost. Avoid delegation for tiny tasks, sequential decisions, edits to the same file, missing authorization, or work where one context must preserve nuanced user intent.

Agent roleUse whenTypical permission scope
iOSSwift/SwiftUI/UIKit implementation or platform reviewAssigned mobile files; tests/build commands
AndroidKotlin/Compose/coroutine implementationAssigned Android files; tests/build commands
BackendAPI, schema, auth, distributed designAssigned service/schema paths; isolated test environment
ArchitectureBoundaries, tradeoffs, ADR reviewRead-only repo/design material unless explicitly assigned edits
SecurityThreat model and control reviewRead-only by default; safe proof-of-concept only in sandbox
TestingUnit/integration/contract/UI test designAssigned tests and local test runner
DevOpsCI/CD, container, deploy/reliability reviewRead-only or staging-only; no production credentials by default
Code reviewCorrectness, security, performance, maintainabilityRead-only diff and test evidence

Give each task: objective, bounded context, owned paths or read-only status, acceptance criteria, prohibited actions, deadline/cancel rule, and output schema. A task response should include status, summary, changed paths, tests/evidence, risks, and unresolved questions. Treat all worker outputs as proposals until verified.

TYPESCRIPT
type TaskContract = {
  taskId: string;
  objective: string;
  contextPackage: { paths: string[]; notes: string[] };
  ownedPaths: string[];
  acceptance: string[];
  permissions: { read: string[]; write: string[]; network: "none" | "allowlisted" };
  budget: { maxInputTokens: number; maxOutputTokens: number; maxToolCalls: number };
};
​
type AgentResult = {
  taskId: string;
  status: "completed" | "blocked" | "failed" | "cancelled";
  summary: string;
  changedPaths: string[];
  evidence: string[];
  risks: string[];
};
​
function validateResult(task: TaskContract, result: AgentResult): void {
  if (task.taskId !== result.taskId) throw new Error("Task identity mismatch");
  if (result.changedPaths.some(path => !task.ownedPaths.includes(path))) {
    throw new Error("Worker changed a path outside its contract");
  }
  if (result.status === "completed" && result.evidence.length === 0) {
    throw new Error("Completed work needs verifiable evidence");
  }
}

The code is an illustrative contract, not an SDK-specific API. Real systems must canonicalize paths, defend against symlinks/path traversal, enforce permissions outside the model, and validate every field at runtime.

Agent lifecycle and recovery

TEXT
received → planned → delegated → running → validating → integrating → reported
                    ↘ awaiting_approval ↗       ↘ retryable_failure
                                                  ↘ blocked / cancelled / failed

Persist task ID, idempotency key, parent/dependency IDs, permission set, model/prompt version, budget reservation, timestamps, result references, and audit events. Define state transitions in code. Set per-tool and overall deadlines; retries are bounded and require idempotent behavior. Cancellation propagates to workers, queues, and long-running tools. Roll back only through an explicit compensating operation or version-control restore plan; do not assume file edits or external actions are transactional.

TYPESCRIPT
type RunState = "planned" | "running" | "awaiting_approval" | "validating"
  | "completed" | "failed" | "cancelled";
​
async function runTask(task: TaskContract, signal: AbortSignal): Promise<AgentResult> {
  await reserveBudget(task.taskId, task.budget);
  await audit({ taskId: task.taskId, event: "started" });
  try {
    const result = await executeWithDeadline(task, signal);
    validateResult(task, result);
    await enforceAcceptanceCriteria(task, result);
    await audit({ taskId: task.taskId, event: "validated" });
    return result;
  } catch (error) {
    await cancelChildren(task.taskId);
    await audit({ taskId: task.taskId, event: "failed", category: classify(error) });
    throw error;
  } finally {
    await reconcileReservedAndActualUsage(task.taskId);
  }
}

reserveBudget, executeWithDeadline, and audit storage are application-owned pseudocode. An implementation needs durable state, race-safe updates, typed errors, and a policy for partial results.

Coordination, DAG execution, and file ownership

Represent dependencies as a directed acyclic graph. Schedule a node only when its prerequisites are validated; cap worker count and resource-intensive builds. Prevent concurrent writers by assigning one owner per path or using isolated worktrees/patches and a serialized integrator. Read-only reviewers can run in parallel against a stable commit. Locking a directory without checking path aliases is not a security boundary.

Example: the iOS worker owns App/Features/Payments/; API worker owns services/payments/; security and architecture reviewers read both contracts; a single integrator owns shared API documentation and resolves interface differences; test worker validates after merge. If the API contract is not stable, make client/server coding dependent on an approved contract rather than having workers guess in parallel.

AI system design and operational quality

An AI gateway should centralize authentication, policy, quota, provider routing, timeouts, retries, redaction, structured output validation, and usage accounting. Store conversation metadata and minimum necessary content with retention and deletion policy. Tool execution should be a separate least-privilege boundary. RAG enforces tenant/resource ACLs before retrieval. Add tracing for request, model call, retrieval, tool, and validator spans; redact prompts, secrets, and sensitive user data by default.

Evaluate offline against versioned golden/adversarial sets and online with opt-in, privacy-reviewed quality signals. Measure task success, groundedness, tool correctness, refusal correctness, p50/p95 latency, retries, human repair, cost, and safety incidents. Canary prompt/model/tool changes; keep rollback. Do not ship on a demo prompt alone.

AI engineering interview hit points

🧠 Say: “I start with a deterministic workflow and add agent autonomy only at points where observations require judgment. Every tool call still passes normal authorization, schema, rate, and audit controls.”

🟨 Tradeoff: A supervisor-worker design improves specialization and parallelism for independent tasks, but costs extra model calls, duplicated context, coordination, and security review. If two tasks share files or depend on unsettled choices, sequence them.

References