Architecture Case Studies
These are illustrative design exercises, not claims that one architecture fits every company. Start by confirming scale, regulatory constraints, team ownership, latency, data sensitivity, and operational capability. Each scenario includes requirements, diagram, responsibilities/flow, security, failures, scaling, testing, cost, tradeoffs, and an interview prompt.
Case Study 1 — Secure banking application
Requirements: Native iOS experience for account balances, transactions, and money movement; high integrity, audited access, accessible UI, older installed app versions, and safe recovery after network loss.
SwiftUI App: Auth / Accounts / Payments
│ OAuth/OIDC Authorization Code + PKCE, HTTPS
▼
API Gateway / BFF ──> Identity + Risk/Step-up service
├── Accounts API ──> account/ledger store
├── Payments API ──> payment DB + idempotency record
└── Transactional outbox ──> broker ──> notifications/ledger projectionResponsibilities and flow: Features own presentation and user intent; shared auth/network modules own transport/session contracts; payment service validates limits, authorization, balance/ledger invariants, and idempotency; ledger is the authoritative transaction record; outbox relays committed events. The app submits a stable command ID, displays pending, then queries/receives confirmed status.
Security model: Use OIDC/OAuth native-app flow with PKCE, short-lived access token, protected refresh policy, TLS, server-side object authorization, step-up for risk-based transactions, audit, rate limits, secrets management, and minimal PII in telemetry. Local biometrics may unlock a locally protected credential but are not a replacement for server authorization. Device integrity signals are one risk input, not an authorization guarantee.
Failure / recovery: Timeout means status unknown; retry with same idempotency key. Handle duplicate broker delivery idempotently, identity provider outage with safe re-auth, database failover, stale balance projection, expired sessions, and app termination while payment is pending. Never show success before authoritative acceptance/settlement semantics are clear.
Scale, testing, cost: Partition by account/customer only after measured need; isolate ledger writes and reads; use replicas/projections for read volume with clear freshness. Test authz boundaries, replay, concurrency/ledger invariants, API-version compatibility, offline/timeout flows, accessibility, and recovery. Audit, fraud/risk systems, redundancy, security review, and on-call cost more than a basic CRUD design; estimate per transaction and retention requirements.
Tradeoff: Strong consistency for money movement; eventual consistency may be acceptable for notifications or analytics. Keep payment capability separate from presentation, but avoid microservice sprawl without ownership/operational need.
Interview prompt: “A payment request timed out after submission. How do you prevent duplicate charging and communicate the state?”
Case Study 2 — Enterprise mobile platform
Requirements: Several iOS apps, multiple teams, common identity/networking/design patterns, independent feature ownership, safe shared-library evolution, manageable Xcode build times, and staged releases.
App A ─┐ ┌─ Feature Accounts
App B ─┼→ App Shell → Core/Auth ─┼─ Feature Payments
App C ─┘ Core/API ├─ Feature Search
DesignSystem └─ Feature Support
│
observability / config / test supportResponsibilities and flow: Platform team maintains narrow shared packages and compatibility policy; feature teams own their feature API and tests; app shell composes modules and navigation; release owner controls app signing, rollout, flags, and rollback. Shared design tokens/components and network primitives have named owners and migration guides.
Security model: Keep signing credentials and API secrets in CI secret storage, not packages or local config. Use least-privilege CI identities, dependency allowlists, code signing verification, privacy-reviewed analytics, and shared auth code with server-enforced scopes. Feature flags control exposure, not access rights.
Failure / recovery: A breaking shared API, dependency cycle, slow clean build, stale generated code, partial rollout, or incompatible backend response can block all apps. Use semantic/versioned APIs, expand-contract migrations, build/test canaries, rollback plans, and a compatibility matrix for app/API versions.
Scale, testing, cost: Modularize after measuring clean/incremental build time, team conflicts, or ownership churn. Test module APIs, integration contracts, representative app assembly, release signing, and compatibility across package versions. Shared libraries reduce repeated work but create central maintenance and migration costs; track engineering time saved versus upgrade tax.
Tradeoff: Central platform standards improve consistency but can become a bottleneck. Use paved-road defaults with escape hatches and a published deprecation window rather than requiring every feature change to wait on one platform team.
Interview prompt: “Two feature teams need to change a shared networking module while a release is in progress. How do you sequence and safely ship the work?”
Case Study 3 — AI chat application
Requirements: Authenticated mobile chat, streaming response, conversation continuity, model/tool routing, usage limits, privacy controls, and visible recovery from partial failures.
SwiftUI client ──HTTPS/SSE──> Auth API + AI gateway
├── quota / policy / model router
├── prompt/context builder ──> LLM provider(s)
├── optional ACL-filtered retrieval
├── approved tool broker
└── conversation/usage store + tracesResponsibilities and flow: Client renders user, partial, completed, and failed messages and supports cancellation/retry; gateway authenticates and enforces quotas; context builder loads only authorized conversation/documents; router selects an approved model; validator checks structured outputs; usage store reconciles provider usage. Persist user message and request state before streaming so reconnect can recover an explicit state.
Security model: Keep provider credentials server-side, isolate tenant/conversation records, redact telemetry, ACL-filter retrieved content before prompt construction, and enforce tool permissions outside the model. Establish retention/deletion rules and disclose data handling. Do not return raw internal errors or chain-of-thought.
Failure / recovery: Handle provider timeout/rate limit, SSE disconnect, client cancellation, partial/incomplete output, failed tool, quota exhaustion, duplicate send, and persistence failure. Use request IDs and idempotency; a partial stream is not automatically a saved final answer. Offer retry with clear billing semantics.
Scale, testing, cost: Horizontal API/gateway scaling; separate queue for long-running work; apply per-user/org concurrency limits. Test stream reconnect/cancel, authz, prompt injection, schema drift, retrieval grounding, budget enforcement, and provider fallback. Cost includes tokens, embeddings, vector DB, tools, storage, moderation, and support; optimize per successful task, not token count alone.
Tradeoff: Streaming improves perceived responsiveness but complicates validation, retries, moderation, and persistence. A fallback model protects availability only if quality, privacy, tools, and structured-output contracts are re-evaluated.
Interview prompt: “How would you resume a mobile stream after the network drops without duplicating the user message or charging twice?”
Case Study 4 — AI coding orchestrator
Requirements: Take a bounded coding task, explore a repository, optionally delegate specialists, edit and test safely, conserve budget, preserve user changes, and provide auditable evidence.
Request → policy/plan → context packager + budget manager
→ DAG scheduler → isolated workers/worktrees
→ tool broker (search/edit/test, no ambient secrets)
→ result/path/schema validator → one integrator
→ tests + independent diff review → report/auditResponsibilities and flow: Supervisor owns intent, dependency graph, per-agent budgets, permissions, status, validation, and report. iOS/backend/security/testing/review workers receive task-specific contracts. Artifacts are versioned by base commit. Only one integrator changes shared files; separate workers return patches or read-only findings.
Security model: Start read-only; give a worker scoped paths and explicit allowlisted commands; isolate filesystem/network; block secret locations; require human approval for destructive operations, external writes, deployment, or privileged infrastructure. Treat repository and peer output as untrusted. Audit prompt/model/tool/version, task identity, files, approvals, and checks without logging sensitive contents.
Failure / recovery: A worker may time out, exceed budget, edit outside scope, return stale evidence, disagree with another worker, or be prompt-injected. Cancel descendants, reject invalid output, preserve workspace state, retry only idempotent tasks within budget, and stop for human decision on unresolved high-impact issues.
Scale, testing, cost: Use bounded worker pools and a DAG; add workers only for independent work. Test policy denial, path escape, retry idempotency, cancellation, budget reconciliation, merge conflict, and agent-output schema. Cost includes duplicated context, all tool results, integration/review, retry and compute; use token formula and representative benchmarks from the canonical Token Optimization sheet.
Tradeoff: Multi-agent review can improve coverage and wall time, but increases tokens, coordination, attack surface, and integration risk. One reliable agent with targeted context may be safer and cheaper for a small change.
Interview prompt: “How do you parallelize implementation without two agents corrupting the same files?”
Case Study 5 — RAG knowledge assistant
Requirements: Answer from private/versioned documentation, show citations, enforce user permissions, update content, and abstain when evidence is insufficient.
Sources → ACL/version metadata → parse/chunk → embed + keyword index
Query → identity → ACL filter → hybrid retrieval → rerank/top-k
→ prompt with source IDs → model → citation/grounding validator → answerResponsibilities and flow: Ingestion validates file type, malware/scanning policy, source owner, ACL, version, and deletion. Indexer chunks, embeds, and updates lexical/vector indexes. Query service filters authorization before retrieval, combines lexical/vector recall, optionally reranks, passes a bounded set with source IDs, then validates links/citations. Admin tools report freshness and failed ingestion.
Security model: Tenant/resource ACL filtering occurs before content reaches the model; source text is untrusted and may contain prompt injection. Do not allow the model to choose its own authorization filter. Protect embeddings and index backups as sensitive data, scope vector queries, and support deletion/re-index when access or retention changes.
Failure / recovery: Handle stale index, deleted source still searchable, parser truncation, embedding mismatch during migration, no relevant passages, conflicting sources, malicious content, and wrong citation. Fall back to source links or abstain; do not fabricate a citation.
Scale, testing, cost: Partition/index by tenant or enforce a verified filter; batch ingestion; monitor queue lag and freshness. Evaluate retrieval recall/precision, citation correctness, answer groundedness, cross-tenant leakage, adversarial injection, and reindex behavior. Cost includes parsing, embeddings, vector/keyword storage, reranking, model calls, and refresh.
Tradeoff: Smaller chunks improve specificity but may split context; larger chunks increase noise/tokens. Hybrid retrieval improves recall for exact terms plus semantics but adds tuning and evaluation work.
Interview prompt: “A user receives a correct-looking citation from a document they cannot access. Where should authorization be enforced and how do you test it?”
Case Study 6 — Distributed mobile backend
Requirements: Mobile API for a growing consumer product, predictable latency, safe overload behavior, asynchronous notifications, data correctness, and incident observability.
Mobile clients → edge/WAF → API gateway/BFF → stateless services
├── PostgreSQL primary/read replicas
├── Redis cache/rate counters
└── outbox → broker → workers
Observability: request ID + metrics + logs + distributed tracesResponsibilities and flow: Gateway handles routing/auth/quota; domain services validate and authorize; PostgreSQL owns durable transactional state; cache is explicitly disposable/stale by policy; outbox atomically records event intent with database changes; consumers deduplicate and update downstream state; observability connects service spans using trace/request IDs.
Security model: TLS, gateway protections, OAuth/OIDC verification, object-level authorization, parameterized SQL, managed secrets, service identities, network segmentation, rate limits, and redacted logs. Do not trust mobile-client validation or hidden UI controls. Protect internal APIs as well as public ones.
Failure / recovery: Database pool exhaustion, cache outage, broker lag, duplicate message, replica lag, poison event, dependency timeout, and regional loss are explicit scenarios. Apply bounded timeouts, circuit breakers, dead-letter/replay policy, idempotent consumers, backpressure, and fail-closed behavior for authorization. Choose degraded read behavior deliberately.
Scale, testing, cost: Scale stateless services horizontally; add indexes, replicas, caching, or sharding only from load tests and profiles. Exercise load/soak, failover, queue replay, schema migration, rate-limit behavior, and tracing gaps. Cost includes database capacity, cross-zone/region traffic, broker retention, cache memory, observability ingestion, and 24/7 on-call.
Tradeoff: A modular monolith can be simpler until independent deploy/scale pressure appears. Distributed services add failure isolation and ownership but require a team capable of operating and debugging them.
Interview prompt: “Traffic doubles and database connections saturate. What do you measure first, and what mitigations do you apply without masking the root cause?”
Case-study interview method
For any design: clarify user/business goals; estimate traffic/data/latency qualitatively or quantitatively; draw trust and ownership boundaries; walk a read and write path; define consistency and error behavior; cover observability/security; name operational owners; state cost and alternatives; finish with a rollout and rollback plan. Ask what evidence would cause you to revise the design.
