Rob Dull/ Tools/ AI ProdOps/ Rebuild retrospective
Build Process · Retrospective

Rebuilding nine tools into one chain.

This page covers the process behind the toolchain: the problem that forced a rebuild, what got built, which tools did the work, how instructions were structured, what got checked in code, how traceability was enforced end to end, where a human stayed in the loop, and what the process taught. The PetHealth case study walks the chain in action; this page covers how it got built.

The starting point

Eight separate AI-assisted tools spanning the product-delivery lifecycle — persona research, journey mapping, feature intake, business-case generation, prioritization, requirements docs, backlog building, sprint planning. Each one worked in isolation. Tracing the actual data flow between them told a different story.

8 → 9
Disconnected tools rebuilt into one real nine-stage chain
2 → 1
Independent scoring implementations, collapsed to a shared source of truth
5
Deterministic guard classes computed in code, never trusted from a model
3
Real bugs found only by human testing, missed entirely by scripted checks
01
The problemDiagnosis

Eight cards sat on the tools index page, each one claiming to hand its output to the next. The prioritization tool wrote no output another tool could consume — it only exported a static file. The backlog tool downstream re-implemented its own prioritization math from scratch, so two independent scoring implementations could silently disagree. Persona data had to be manually copy-pasted into the journey-mapping tool. Epic-level and feature-level planning were conflated in the shared schema, so prioritization decisions and delivery-team sizing decisions were happening on the same object when they needed to be separable.

The work behind "automated" was several people's worth of manual copy-paste, run by hand behind a clean UI.

02
What got builtRebuild

A restructured, nine-stage chain, deployed live: Persona → Journey Mapping → Initiative Intake → Business Cases → Epic Prioritizer → Roadmap & Milestones → Business Documents (the BRD) → Backlog Builder → Sprint Planner, ending in a Jira-importable CSV. Every stage's output became the next stage's real input, through the tools' same-origin handoffs.

The demo runs an insurance-claims initiative through the whole chain with both a front-stage persona and a backstage persona, deliberately, so the chain has to handle customer-facing and infrastructure work correctly — exactly where the original tool broke down: backstage work kept getting mis-scoped as if it were customer-facing.

03
Tools usedStack
The build leaned on a deliberate set of tools, each doing a specific job.
ClaudeMultiple model tiers by phase — a faster model to scaffold a stage, a stronger model for a dedicated adversarial review pass.
Claude CodeThe agentic build environment: multi-step builds, live browser verification, and file-level review in one loop.
React + TypeScript monorepoOne shared toolkit package for data types and scoring math, so the duplicate-implementation bug can't recur by construction.
MCP-style evidence layerFixture data standing in for real connectors, so every generated artifact has something real to cite.
04
Structured instructionsPrompting

Nothing in the chain runs on free-text handoffs. Every stage has a typed input/output contract defined once in a shared package, and the model fills in that fixed schema.

Two instruction patterns did the most work. Cite only what's present: every generative step references evidence by a [REF] id and may only use ids actually included in its input — anything it lacks evidence for is marked ASSUMED. Keep independently-varying concepts decoupled: an early version conflated "is this epic customer-facing or infrastructure" with "does its business case rest on a customer metric or an engineering one" — two genuinely independent questions collapsed into one field, producing wrong framing downstream. Splitting them into two fields, classified independently, was verified with a controlled test: two epics with identical infrastructure classification produced correctly divergent business-case framing.

05
Quality checksGuardrails
The rule running through the whole build: the model never gets to be the source of truth for anything checkable. Wherever an output could be computed or verified deterministically, it is — in code, downstream of the model, every time.
★ Enforced in code
  • Prioritization scores recomputed from structured inputs in code
  • Requirement coverage checked: every business requirement must be addressed by a downstream feature before a run proceeds
  • Dependency ordering enforced: a plan that schedules a story before something it depends on is rejected, checked against every export path
  • Output shape defended everywhere: one missing field from a model response can't take down a page
06
TraceabilityProvenance
Every artifact downstream carries forward the ids of what produced it. Code enforces that nothing references an id that doesn't actually exist upstream.
BR-04→ FEAT-02→ ST-07→ ST-09→ Jira: Blocks
A feature cites the business-requirement ids it satisfies; a story cites its parent feature and the other stories it depends on; the final export's dependency links are read directly off that same story graph. An exec summary generated at one stage seeds directly into the next stage's document.
07
Human reviewJudgment

Two tiers. First, a second AI pass on complex phases, framed as adversarial review, hunting for bugs. Second, and more important: every phase was live-tested with real inputs, separate from the fixture data used for automated verification.

Scripted verification proved every code path executed. Hand-authored test fixtures are too clean to surface real bugs.

★ Caught only by a human running the real workflow
  • A paste box silently accepted the wrong shape of data because two nearly-identical text boxes sat next to each other on the same screen
  • A rule displayed a correct warning on screen but silently vanished from the exported file
  • Generated documentation described a feature the prompt explicitly said shouldn't appear
✋ Human gate: every phase reviewed with real inputs before it counted as done
08
What the process taughtRetrospective
01

Scripted and human testing catch different bugs

Skipping the human pass on anything with a paste or free-form intake step is exactly where the real bugs hide.

02

Never trust a model's own math or ordering

Recompute deterministically downstream, every time — even after it gets something right once.

03

Design for visible failure

A system that marks content ASSUMED or refuses below a threshold is more trustworthy than one that always sounds confident.

04

Test the connections end to end

Each individual tool working in isolation is exactly the gap that caused the original failure and started this rebuild.

05

Decouple concepts that vary independently

Folding two real distinctions into one field for simplicity produces silently wrong downstream framing that's hard to trace back to its cause.

What this process demonstrates

  • Every stage was tested against the next stage's real handoff, end to end
  • Deterministic guards ran downstream of every generative step
  • Traceability held across the full chain: an id from an early stage is still checkable at the last one
  • Live human testing with real inputs found real bugs that scripted, fixture-based verification missed

Scope

  • This is a working prototype. Scope disclosures are on the case study and technical notes pages
  • Human judgment applies at each gate, alongside the guardrails
  • Model behavior varies run to run; the computed checks are the repeatable part
  • The second-model adversarial review ran on select complex phases