01Progressive data service
A versioned key-value store that gains transactions, rollback, expiry and concurrency in stages. Include tests and a decision log for every refactor.
Proves: fundamentals under changing requirements
02Reliable webhook platform
Multi-tenant delivery with signing, idempotency, retry policies, dead-letter queues, observability and a written SLO. Break it deliberately and publish the postmortem.
Proves: production systems judgment
03Auditable agent runtime
A typed model–context–tool loop with durable sessions, disposable sandboxes, explicit context budgets, approval boundaries, trace replay and crash recovery that never repeats a confirmed side effect.
Proves: harness and state engineering
04Regulated-domain agent pilot
A narrow workflow with a domain expert, 60-task eval suite, tenant permissions, citations, audit trail, prompt-injection tests, fallback, feature flag and adoption metric. Start read-only; graduate one reversible action.
Proves: AI reliability and earned trust
05Agent incident game day
Combine poisoned retrieval, an unauthorized tool attempt and provider throttling. Measure detection, containment and recovery; publish the trace, blast radius, counterfactual controls and regression task.
Proves: operational ownership under failure
06Outcome case study and handoff
A reproducible package: repository, architecture record, eval suite, threat model, runbook, pilot metrics, incident report and narrative covering baseline, rejected options, impact, failures and reusable product feedback.
Proves: end-to-end FDE ownership and transfer