Rob Dull/ Tools/ AI ProdOps/ Case study
End-to-End Case Study · Real Captured Runs

One initiative, nine tools, zero copy-paste.

This page traces a single initiative through the entire chain: from an enterprise persona to a Jira-ready sprint plan. Every handoff below happened through the tools' real same-origin handoffs, one click each. Every step links to the captured run itself, so you are judging real output, not a slideware summary of it. The human work in this chain is the review at each gate; the drafting is the AI's.

Evidence disclosure

PetHealth is fictional. The company, people, source records, metrics, approvals, and operating environment in this example were created for demonstration. The displayed artifacts use the structures produced by the tools and are intended for evaluation.

The example contains four kinds of information: scenario facts (fictional inputs intentionally supplied to keep the test coherent, such as the 22% support-contact baseline), fixture evidence (fictional records with source-like identifiers used to test provenance and propagation), model-generated content (draft requirements, assumptions, estimates, NFRs, risks, and narrative), and deterministic results (selected calculations and structural conditions recomputed in code).

A fixture citation shows that a generated claim was associated with a supplied record. It does not make the fictional claim independently true. The review gates shown in this case are modeled checkpoints, not records of approvals by real PetHealth employees.

The scenario

PetHealth is a fictional mid-size pet insurer with a real-shaped problem: members don't trust the claims process. Claim-status questions alone account for 22% of all support contacts, members describe the period after submitting a claim as a black hole, and leakage in adjudication quietly erodes margin. The company context, evidence set, and competing priorities are held constant across every tool, which is what makes the chain a fair test: each tool inherits its input from the one before, not from a hand-tuned prompt.

The names are fictional. The runs are not: each artifact below is captured output from the live tool.

1 → 5
One funded initiative, decomposed into five epics with framed assumptions
8
Business requirements in the BRD, each traceable downstream
7 / 15
MoSCoW'd features and ordered stories in the extended backlog
8 of 8
BR coverage in the backlog: computed by code, not claimed by the model
01
Enterprise PersonaDiscover
A rough description of a PetHealth member becomes a full enterprise persona, detailed enough to anchor everything downstream.
★ What came out

Marcus Chen, software engineer and PetHealth member. His quote sets the tone for the whole initiative: "I just want to know it's working." Until he needs it. The card carries his Jobs to Be Done statement, tech comfort across channels, key interactions with friction badges, and the quantified organizational impact of his unmet job.

HANDOFF → the persona seeds the journey map directly.
What to evaluate: Does the persona clarify the job and operating context, or merely add plausible detail? Which attributes would require real research?
Open the tool: it is its own demo ↗ ✋ Human gate: PM ratifies the persona before mapping
02
Journey Map BuilderMap
Marcus's journey across pre-enrollment, account management, claims, and renewal, with a draggable sentiment curve and row-level pain points. The sentiment trough lands exactly where the support data says it should: the silence after a claim is submitted.
★ What came out

A stage-by-stage map whose pain clusters become opportunity seeds: structured, exportable statements of where the experience breaks and what it costs.

HANDOFF → opportunity seeds feed initiative intake.
What to evaluate: Is the pain point actually supported by evidence, and does the map distinguish observation from interpretation?
03
Initiative IntakePortfolio tradeoff
Four sequential AI calls weigh Claims Transparency & Leakage Reduction against two competing initiatives: Renewal & Pricing Experience Revamp and the Vet-Clinic Partner Portal, using cited strategy, budget, and demand evidence. The funded initiative is then decomposed into five epics, split between front-stage business work and backstage architecture runway.
★ What came out
  • EPIC-01 Auto-Adjudication Engine Modernization
  • EPIC-02 Reimbursement Payout Upgrade
  • EPIC-03 Real-Time Claim Status Tracking
  • EPIC-04 Vet-Clinic Direct Submission Integration
  • EPIC-05 Fraud & Leakage Analytics

Each epic carries its riskiest Demand, Feasibility, and Viability assumptions and a benefits-realization plan naming the measurement instrumentation that has to exist. The headline benefit target: claim-status share of support contacts falls from 22% to ≤12%.

HANDOFF → the five epics hand off to Business Cases, one case per epic, carrying their D/F/V risk surface with them.
What to evaluate: Are the alternatives and epic boundaries credible? Which assumptions are useful risk framing, and which are invented context?
View the captured run ↗ ✋ Human gate: PM + stakeholders sign off on the intake
04
Business CasesJustify
One case per epic: value, cost, strategic alignment, cost of inaction, and a rough order-of-magnitude effort estimate. That estimate is not decoration: it anchors the job-size input the prioritizer uses next, so the ranking and the funding conversation share one number.
★ What came out

Five business cases in a standard format, including the Real-Time Claim Status Tracking case that the rest of this page follows.

HANDOFF → the cases become the prioritizer's candidate records.
What to evaluate: Would the case support an investment conversation after correction, or does its format overstate the quality of its evidence?
View the example gallery (Claim Status tab) ↗ ✋ Human gate: sponsor sign-off per case
05
Epic PrioritizerRank
All five epics scored on WSJF and RICE components in a single pass, customer-facing features and infrastructure work on the same board. The weights are editable and the board recomputes live: raise risk-reduction and an architecture epic can take the top slot with no re-run. This is the page to show whoever in your org still ranks epics in a spreadsheet fight.
★ What came out

A defensible ranking with an explicit funding line. Real-Time Claim Status Tracking clears it, and becomes the epic the delivery half of the chain executes.

HANDOFF → the ranking plus funding line feed the roadmap.
What to evaluate: Are the inputs defensible? Does one board improve the funding conversation, or create false comparability?
06
Roadmap & MilestonesSchedule
Every prioritized epic sequenced into release windows across up to 18 months, funded epics first, the rest as capacity frees. Dependencies stay deliberately light: named approvals or genuine architecture-epic runway only, so the roadmap stays a communication tool instead of a Gantt chart.
★ What came out

The Roadmap: one of the four primary outputs program stakeholders care most about.

HANDOFF → the roadmap gates which epic gets a BRD.
What to evaluate: Does the roadmap preserve priorities and dependencies without implying knowledge of capacity that was never supplied?
07
Business DocumentsSpec
A standard-template BRD for the funded epic: scope, prioritized business requirements, and high-level NFRs, plus the executive package (summary, objectives, stakeholders, cost-benefit, sign-off). One structural guardrail worth noticing: the tool refuses to write a BRD for an epic below the funding line. Process discipline enforced in code, not in a wiki page nobody reads.
★ What came out

The BRD for Real-Time Claim Status Tracking with eight business requirements (BR-01 through BR-08), an assumptions section, and the exec sign-off package.

HANDOFF → the BRD is the single primary input to the backlog.
What to evaluate: Which sections are reviewable, which need correction, and which should be omitted until evidence exists?
View the captured run: BRD with Assumptions ↗ ✋ Human gate: architecture, security, and business review
08
Backlog BuilderPlan
The BRD becomes the full delivery-doc suite: features, stories, solution outline, security and UAT plan, and dev and user docs. The part that matters for trust: requirement coverage is computed, not generated. The tool checks, in code, that every business requirement is covered by at least one feature, and flags sizing variance against the earlier estimate instead of quietly smoothing it over.
★ What came out
  • 7 MoSCoW'd features (FEAT-01 to FEAT-07) with feature-level NFRs and refined sizing
  • 15 ordered stories with BDD acceptance criteria, points, and dependsOn sequencing
  • Computed coverage: 8 of 8 business requirements covered, with the refined-versus-rough sizing variance flagged
HANDOFF → the extended backlog carries its dependency graph into sprint planning.
What to evaluate: Does identifier coverage reflect meaningful satisfaction of the requirement? Would engineering consider the stories sufficiently grounded and testable?
09
Sprint Planner & Jira ExportDeliver
Two AI calls turn the extended backlog into a delivery plan: MVP scope from the Must-priority features, a sprint-by-sprint allocation, baseline and worst-case estimates with the drivers that widen the gap, and a RAID log. Sprint loads and dependsOn ordering are recomputed in code: an over-capacity sprint or a story scheduled before its dependency raises a visible warning wherever the plan is consumed.
★ What came out

The delivery plan, and the chain's outbound handoff: a Jira bulk-import CSV with Summary, Issue Type, Epic Link, Priority, Story Points, Sprint, Acceptance Criteria, and Issue Links. The chain ends inside your delivery system, not in a document nobody imports.

What to evaluate: Are the estimates credible? What work or constraints are missing? Would the CSV import cleanly and preserve the intended relationships?
View the captured run ↗ ✋ Human gate: delivery lead owns the plan before import

What this run demonstrates

  • The handoffs are real: each tool consumed the previous tool's primary output, unedited, through same-origin handoffs
  • Traceability survives the whole chain: a support-contact statistic from the scenario context shows up as a benefits target in intake, a requirement in the BRD, and covered stories in the backlog
  • Business and architecture work compete on one board instead of in separate meetings
  • The guardrails hold: funding-line refusal, computed coverage, and sprint-load and dependency warnings all fired where they should

What it does not claim

  • PetHealth is fictional: this shows the chain working end to end, not a customer deployment
  • A human reviewed at each gate: the chain removes drafting toil, not judgment
  • Model output varies run to run; the computed checks are the repeatable part
  • The MCP data sources shown in the network diagram were not live for this run; the evidence set was provided as context

Known corrections and open questions

This section should be updated as external evaluations occur. Current issues visible in the worked example: