This page covers the process behind the toolchain: the problem that forced a rebuild, what got built, which tools did the work, how instructions were structured, what got checked in code, how traceability was enforced end to end, where a human stayed in the loop, and what the process taught. The PetHealth case study walks the chain in action; this page covers how it got built.
Eight separate AI-assisted tools spanning the product-delivery lifecycle — persona research, journey mapping, feature intake, business-case generation, prioritization, requirements docs, backlog building, sprint planning. Each one worked in isolation. Tracing the actual data flow between them told a different story.
Eight cards sat on the tools index page, each one claiming to hand its output to the next. The prioritization tool wrote no output another tool could consume — it only exported a static file. The backlog tool downstream re-implemented its own prioritization math from scratch, so two independent scoring implementations could silently disagree. Persona data had to be manually copy-pasted into the journey-mapping tool. Epic-level and feature-level planning were conflated in the shared schema, so prioritization decisions and delivery-team sizing decisions were happening on the same object when they needed to be separable.
The work behind "automated" was several people's worth of manual copy-paste, run by hand behind a clean UI.
A restructured, nine-stage chain, deployed live: Persona → Journey Mapping → Initiative Intake → Business Cases → Epic Prioritizer → Roadmap & Milestones → Business Documents (the BRD) → Backlog Builder → Sprint Planner, ending in a Jira-importable CSV. Every stage's output became the next stage's real input, through the tools' same-origin handoffs.
The demo runs an insurance-claims initiative through the whole chain with both a front-stage persona and a backstage persona, deliberately, so the chain has to handle customer-facing and infrastructure work correctly — exactly where the original tool broke down: backstage work kept getting mis-scoped as if it were customer-facing.
Nothing in the chain runs on free-text handoffs. Every stage has a typed input/output contract defined once in a shared package, and the model fills in that fixed schema.
Two instruction patterns did the most work. Cite only what's present: every generative step references evidence by a [REF] id and may only use ids actually included in its input — anything it lacks evidence for is marked ASSUMED. Keep independently-varying concepts decoupled: an early version conflated "is this epic customer-facing or infrastructure" with "does its business case rest on a customer metric or an engineering one" — two genuinely independent questions collapsed into one field, producing wrong framing downstream. Splitting them into two fields, classified independently, was verified with a controlled test: two epics with identical infrastructure classification produced correctly divergent business-case framing.
Two tiers. First, a second AI pass on complex phases, framed as adversarial review, hunting for bugs. Second, and more important: every phase was live-tested with real inputs, separate from the fixture data used for automated verification.
Scripted verification proved every code path executed. Hand-authored test fixtures are too clean to surface real bugs.
Skipping the human pass on anything with a paste or free-form intake step is exactly where the real bugs hide.
Recompute deterministically downstream, every time — even after it gets something right once.
A system that marks content ASSUMED or refuses below a threshold is more trustworthy than one that always sounds confident.
Each individual tool working in isolation is exactly the gap that caused the original failure and started this rebuild.
Folding two real distinctions into one field for simplicity produces silently wrong downstream framing that's hard to trace back to its cause.
This page is the process behind the chain. To see the chain itself running against one real fictional initiative, walk the captured run. To see the actual code — data contracts, guardrails, and the trust boundary around API keys — read the engineering notes.