Afflatus
SURVEILLANCE FIELD MANUAL · ARCHIVE 52-FDE

A 52-week evidence route for turning frontier AI into governed, adopted production systems.

Forward deployed
engineer.

Advance transmission ↓
00 · ORIENTATION INTELLIGENCE

FDE is not one job. It is the control plane for enterprise AI change.

My read as of 2 August 2026: the title is in a land-grab, but the function is durable. Model access is commoditising; the scarce work is choosing a valuable workflow, crossing security and data boundaries, proving reliability, earning adoption, and feeding field truth back into product.

ANALYST NOTE / CLAIMS ARE DATED · THIRD-PARTY SAMPLES ARE LABELLED

THESIS / 01

Deployment, not intelligence, is the bottleneck.

The 2026 capital moves by OpenAI and AWS are stronger signals than title hype: both are funding organisations that work inside customer data, controls and workflows. The economic unit is no longer an API call; it is a changed business process with measurable adoption.

THESIS / 02

The durable skill is translation with consequences.

An FDE translates tacit domain rules into software, business risk into permissions and evals, and production evidence into product decisions. Prompt syntax is a small surface. Judgment across engineering, operations, security and incentives is the real craft.

THESIS / 03

The best FDE makes the customer less dependent.

A deployment that only one heroic engineer can maintain is consulting debt. The winning output is a live system plus an owned evaluation set, observable operations, reusable components, a decision trail and a customer team able to run it. AWS now states self-sufficiency as an explicit end condition.

THESIS / 04

This is senior work even when the title is not senior.

The environment—not the badge—sets the level: unfamiliar code, political ambiguity, production risk and executive visibility. A zero-to-FDE plan should build a production engineer first, then add field discovery and ownership. One year can create evidence; it cannot honestly promise a frontier-lab offer.

FOUR MARKETS HIDING UNDER ONE TITLE

Choose the operating model before choosing the employer.

01

Frontier lab

Own strategic deployments; turn model failures and workflow evidence into research and product feedback.

02

Cloud or data platform

Build governed agent systems on a platform; codify delivery into repeatable harnesses and partner capability.

03

Integrator or consultancy

Combine engineering with industry redesign and change management across large estates; delivery breadth is the product.

04

Vertical AI company

Live inside one domain, discover the repeatable core, and turn bespoke work into product. Highest learning rate; highest services-trap risk.

BACKGROUND YOU MUST KNOW

The model is one dependency in a much larger production system.

  1. 01

    Enterprise boundariesIAM, SSO, RBAC, data residency, audit, procurement and legal review.

  2. 02

    Agent reliabilityTask-level evals, least privilege, human escalation, monitoring, controlled rollout and rollback.

  3. 03

    Workflow economicsBaseline, counterfactual, adoption, cost per completed outcome and failure cost—not demo accuracy.

  4. 04

    Domain discoveryObserve exception paths, incentives, tacit knowledge, decision rights and who absorbs errors.

  5. 05

    Production engineeringAPIs, databases, queues, distributed failure, observability, cost, latency and incident response.

  6. 06

    Adoption and transferTraining, operating procedures, documentation, champions and a handoff that survives your departure.

01 · TEXTBOOK AUDIT / AGENT ENGINEERING

Do not read 318 pages in order. Turn the book into a production loop.

Li Bojie's v1.4 textbook is already dated 1 August 2026 and is unusually current. This course therefore preserves its engineering spine, challenges its volatile claims, and reorders its experiments around what an FDE must prove in a customer environment.

KEEP / 01

The durable spine

Agent = model + managed context + tools. Keep the chapters on context, memory and retrieval, tool contracts, event-driven execution, evaluation and learning from traces. These are architectural constraints, not product trivia.

REFRAME / 02

The claims that need tests

RAG is not memory; a larger context is not better context; an LLM judge is not ground truth; more agents are not automatically more capable. Treat every pattern as a hypothesis and compare it with a simpler baseline under matched tools, budgets and environments.

ADD / 03

The missing customer boundary

Add identity, authorization, data residency, sandbox isolation, durable sessions, incident response, cost per completed outcome, rollout and rollback, adoption and handoff. A technically impressive agent that cannot cross these gates is not a deployed system.

ELECTIVE, NOT COREPost-training, realtime voice, computer use, robotics and agent societies remain valuable specialisations. Enter them after the production-agent gate, or earlier only when your target deployment demands them.

THE ROLE, WITHOUT THE HYPE

Not a consultant who stops at slides. Not a researcher who stops at a benchmark.

An FDE owns the full loop: discover the workflow, scope the risk, build the system, measure adoption, and turn field lessons into reusable product capability.

  1. 01Discover the costly decision
  2. 02Prototype the smallest proof
  3. 03Engineer for failure
  4. 04Earn adoption and trust
  5. 05Measure value, feed product
YOUR LEARNING PATH

One task. One piece of evidence.

Follow the 36 packets in week order. Prerequisites are recommendations; every packet remains open to explore.

Progress and evidence stay in this browser, without an account. Export a copy before clearing browser data. Completion is your own assessment, not an independent review.

Browse all 36 resources in the map. Progress tracking needs JavaScript.

02 · THE FIELD MAP

Six hard questions. Thirty-six transmissions. One page.

The route develops into an open evidence field, inspired by Anthropic's Path to Hope: every article or assignment orbits a hard question instead of sitting inside a dashboard. Fine wires reveal its home question; brighter lines expose ideas that transfer across clusters. Select any cover for theory, a build specification, a fault injection and an evidence gate.

SCROLL TO DEVELOP THE FILM ↓
OPEN FIELD / 36 TRANSMISSIONS / 16 STRONG LINKS
Projector

The articles are spread across the entire projector field. Hairlines lead each cover back to its question; focus a cover to reveal its strongest cross-question transfer. Books drift, repel nearby evidence and never trap page scrolling. Drag only after zooming.

WEEK 01 · TRANSMISSION

Course packet

THEORY

BUILD

BREAK

EVIDENCE GATE

Open transmission ↗
03 · 52-WEEK SEQUENCE

A calendar with gates, not deadlines.

Advance only when the evidence gate is met. Repeat a phase when the artifact is weak; skip material you can already prove. Budget: 10–12 focused hours each week, split between theory, construction, fault injection and evidence.

  1. W01–02

    Orientation

    Write a one-page role thesis and interview two practising builders. Gate: explain where an FDE creates value and where the role may disappear.

  2. W03–08

    Software foundations

    Build a hash table, LRU cache, task queue and versioned key-value service without AI writing the first draft. Gate: tests, profiling, bounded memory, async I/O and an architecture decision record.

  3. W09–14

    Production runtime

    Turn the service into a multi-tenant webhook platform. Gate: transactions, outbox, idempotency, retries, dead-letter queue, written SLO, traces, load test and rollback.

  4. W15–24

    Recoverable agent runtime

    Implement a model–context–tool loop under 100 lines, then separate session, harness and sandbox. Gate: typed tool contracts, context-budget experiment, append-only replay, snapshots and crash recovery without repeating side effects.

  5. W25–32

    Memory, events and security

    Build hybrid retrieval, execution-state memory and an interruptible event loop; expose one specialist through A2A. Gate: freshness and authorization evals, cancellation and resume, prompt-injection red team, threat model and verified containment.

  6. W33–40

    Evaluation and orchestration

    Create 60 production-grounded tasks and measure repeated trials under resource, prompt and tool faults. Gate: confidence intervals, grader disagreement, reliability surface, and a matched comparison proving whether multi-agent complexity earns its cost.

  7. W41–52

    Governed customer deployment

    Shadow five users, model cost per completed outcome and ship one narrow pilot to 5–10 people. Gate: least privilege, adoption and override metrics, rollback drill, incident postmortem, customer handoff and a reproducible outcome case study.

04 · DEGREE, COURSE OR FIELD?

Choose the missing signal, not the most prestigious label.

PATH A

Portfolio-first

Best when you already ship software. Use the 52-week sequence, domain fieldwork and public case studies. LinkedIn reports an 8.2× larger AI talent pipeline when companies focus on skills over degrees or titles.

Signal: shipped outcomes
PATH B

BSc CS / Software Engineering

Worth considering when systems, algorithms and mathematical foundations are genuinely missing — especially if structured internships and peers will accelerate you. Do not expect the degree alone to demonstrate customer judgment.

Signal: durable foundations
PATH C

MSc CS / AI

Useful for deeper ML systems, research collaboration or a geography change. Optional for FDE: current role descriptions emphasize production experience, ambiguity and communication more directly than a graduate credential.

Signal: technical depth
PATH D

Domain credential

High leverage in health, finance, public sector, energy or defence. A short domain qualification plus supervised fieldwork can make you safer and more useful than another generic AI certificate.

Signal: earned trust

Before committing, ask yourself

  • Do I enjoy spending days inside another team's messy workflow?
  • Can I travel, handle executive pressure and still make calm technical decisions?
  • Which domain do I care enough about to learn its language, regulation and failure costs?
  • Am I willing to say 'an LLM is the wrong tool' after weeks of discovery?
05 · THE SIX PORTFOLIO PROOFS

Build a ladder of evidence, not a pile of demos.

01

Progressive data service

A versioned key-value store that gains transactions, rollback, expiry and concurrency in stages. Include tests and a decision log for every refactor.

Proves: fundamentals under changing requirements
02

Reliable webhook platform

Multi-tenant delivery with signing, idempotency, retry policies, dead-letter queues, observability and a written SLO. Break it deliberately and publish the postmortem.

Proves: production systems judgment
03

Auditable agent runtime

A typed model–context–tool loop with durable sessions, disposable sandboxes, explicit context budgets, approval boundaries, trace replay and crash recovery that never repeats a confirmed side effect.

Proves: harness and state engineering
04

Regulated-domain agent pilot

A narrow workflow with a domain expert, 60-task eval suite, tenant permissions, citations, audit trail, prompt-injection tests, fallback, feature flag and adoption metric. Start read-only; graduate one reversible action.

Proves: AI reliability and earned trust
05

Agent incident game day

Combine poisoned retrieval, an unauthorized tool attempt and provider throttling. Measure detection, containment and recovery; publish the trace, blast radius, counterfactual controls and regression task.

Proves: operational ownership under failure
06

Outcome case study and handoff

A reproducible package: repository, architecture record, eval suite, threat model, runbook, pilot metrics, incident report and narrative covering baseline, rejected options, impact, failures and reusable product feedback.

Proves: end-to-end FDE ownership and transfer
06 · INTERVIEW REHEARSAL

Practice the work in the shape it will arrive.

60 MIN

Progressive coding

Start correct, then absorb four requirement changes. Narrate state, memory, concurrency and rollback trade-offs. No clever rewrite at the end.

60 MIN

Customer solution design

Clarify workflow, users, risk, volume and success before drawing architecture. Surface failure modes and a non-LLM alternative without being prompted.

45 MIN

Artifact deep dive

Defend one project line by line: what you knew, what you assumed, what failed, what the data changed, and what you would remove now.

45 MIN

Values under pressure

Prepare real stories about pushing back, changing your mind, negative feedback and an ethical conflict. Name the emotion and the concrete trade-off; avoid rehearsed virtue.

07 · WEEKLY EVIDENCE LOOP

Your prompt and artifact become next week's curriculum.

This worksheet never leaves your device. Capture one prompt, one artifact and one decision; score only visible evidence, then carry the generated brief into your next Codex conversation.

Score only what you can show

FIELD SCORE0/ 100

Add evidence before scoring.

Anti-gaming rule A score without a link, test, transcript, metric or decision record is capped at 4/10.

Adaptation rule Next week spends 50% of deliberate practice on the weakest evidence dimension, 30% on the active phase gate, and 20% on curiosity.