Rob Dull/ Tools/ AI ProdOps/ Limited Evaluation Beta runbook
Limited Evaluation Beta · Operator Guide

The beta runbook: one initiative, one sitting.

You've read the pitch and decided to try it. This page is the operator guide: what to have ready, the suggested 60-to-90-minute path through the chain, what to pay attention to at each gate, and how to send your run back when you're done. Nothing here requires reading anything else first.

Use fictional, public, or sanitized information only

This is a single-user proof of concept, not a secure or production system. Do not enter confidential strategy, customer personal data, proprietary financials, source code, security architecture, or regulated or export-controlled information. A realistically disguised initiative, with names and numbers changed, works just as well as the real thing for evaluating the tool.

What to bring

Two tiers. Tier 1 is enough to run the chain end to end and judge the mechanics. Tier 2 is what turns the run into a genuine test of output quality, because a few of the steps below quietly invent plausible-sounding detail when you don't give them anything real to work with. Bring what your time allows; more real input produces a more honest verdict.

Tier 1 · Minimum

~5 min prep
01
One initiative, one paragraph
A real problem from your org, or a realistically disguised one: what's broken, who it hurts, and one number that shows it matters.
02
An Anthropic API key (optional)
From console.anthropic.com. No key? Walk the same path through the captured sample runs instead; you'll evaluate output quality, just not on your own initiative.
03
One browser, kept open
The chain hands work between tools inside your browser's storage. Stay in one browser for the session. You can download your whole run as a file at any point (see below), so nothing is lost if you stop early.

Tier 2 · Recommended for a genuinely evaluative run

+15–20 min prep
Without the items below, Initiative Intake (step 03) will invent plausible-sounding competing initiatives and a budget assumption to fill the gap, because that's what the prompt is designed to do when nothing real is given. That's a reasonable fallback for a five-minute look, but it isn't a fair test of the quality of the tool's draft and structure. Bring the real thing and you're testing the tool; skip it and you're mostly testing the model's imagination.
  • Two named competing initiatives, each with a one-line reason they lose the capacity to yours. This is what actually exercises the portfolio-tradeoff step instead of letting it improvise.
  • A real quote, stat, or budget line, pasted below your paragraph as an evidence line so the model cites it with provenance instead of assuming it. The format:
    [REF-SUPPORT-01 support-ticket] "Members message support asking where their claim is; we can't tell them anything past 'still processing.'"
    Any short id and kind works, e.g. [REF-BUDGET-Q3 financial-tracking] or [REF-NPS-091 external-comms]. This format isn't shown anywhere in the tool itself, the visible evidence chips are fixture demo data, so bringing your own only works if you know to paste it this way. A citation like this associates a claim with the record you supplied; it does not verify that the record itself is accurate.
  • A one-line budget ceiling or OKR the initiative should ladder to, so the tradeoff rationale grounds in something real.
  • If you'll reach Backlog Builder (step 08): an excerpt of your org's actual DoR/DoD, a style-guide snippet, and one line on your current-state architecture. All three are real optional fields on that tool and visibly change the NFRs and solution outline from generic to grounded.
  • If you're starting from the top (steps 01–02): real persona or interview notes for the Journey Map input rather than a cold start. Its own placeholder text says it plainly: the more specific, the better the suggestions.

Intake to sprint plan, seven stops

This is the recommended beta path. It starts at Initiative Intake (step 03) rather than the very top of the chain: personas and journey maps are worth a look later, but the intake-to-Jira stretch is where the system earns or loses your trust. Each stop names what to watch for, because where the output falls short of your org's bar is exactly the feedback this beta exists to collect.

03

Initiative Intake · ~15 min

Paste your one-paragraph initiative and run the four-step chain: portfolio tradeoff, epic decomposition, riskiest assumptions, benefits realization. Review each step before running the next; edit anything that's wrong before it flows downstream.

Watch for Would the Demand / Feasibility / Viability assumptions survive your stakeholders? Are the epics it framed genuinely yours, or generic? What did the model introduce that was not in your input?
04

Business Cases · ~10 min

One click on a framed epic pre-fills the case form. Generate a case for one or two epics, not all of them; the point is judging the quality bar, not exhaustive coverage.

Watch for Would you put this case in front of your sponsor? Is the rough effort estimate credible enough to anchor prioritization? What did the model introduce that was not in your input?
05

Epic Prioritizer · ~10 min

Score your cases on WSJF and RICE in one pass, then drag the weights around. The board recomputes live, including the funding line.

Watch for Does the ranking hold up when you raise risk-reduction? Would this board survive your loudest stakeholder? What did the model introduce that was not in your input?
06

Roadmap & Milestones · ~5 min

One click carries the ranking here; the funded epics sequence into release windows with a deliberately light dependency list.

Watch for Are the dependencies real blockers or invented ceremony? What did the model introduce that was not in your input?
07

Business Documents · ~10 min

Generate the BRD for your funded epic. Note that the tool refuses epics below the funding line; that's a guardrail, not a bug.

Watch for The requirements core and the executive package: which sections would your org keep, cut, or fight about? Flag any NFR threshold, financial figure, stakeholder name, or sign-off that reads as more certain or more approved than the evidence supports. What did the model introduce that was not in your input?
08

Backlog Builder · ~15 min

The BRD becomes MoSCoW'd features and ordered stories with BDD acceptance criteria. Requirement coverage is computed against the BRD, not claimed.

Watch for Would your engineers pick up these stories? Are the acceptance criteria testable or vague? What did the model introduce that was not in your input?
09

Sprint Planner · ~10 min

MVP scope, a capacity-and-dependency-checked sprint plan, baseline and worst-case estimates, a RAID log, and the Jira bulk-import CSV that ends the chain.

Watch for Do the sprint loads and the worst-case gap reflect how delivery actually goes where you work? What did the model introduce that was not in your input?
+

If you have another 20 minutes

Start from the top instead: generate a persona (01), build a journey map from it (02), and let its pain clusters seed the intake. That's the full chain, research to Jira.

Where your initiative goes

Your browser to Anthropic and back; nothing else

Your key lives in session memory and clears when the tab closes. Your inputs travel from your browser to Anthropic's API and back; there is no server of mine in the path, no account, and no database. All saved artifacts live in your browser's localStorage. If your initiative is sensitive, disguise the names and numbers; the chain works just as well on a realistic fiction. The full statement, including what is and isn't verifiable, is on the technical notes page.

The run file: your save button and your feedback channel

Download early, download often

At the bottom of every tool there's a Run file row with Download and Load buttons. Download snapshots everything the chain has saved, from any tool, into one JSON file. Load restores it in any browser, on any machine. Download whenever you finish a stage; if a cache clear or a browser switch would otherwise lose your session, the file gets it all back.

It's also how you can send me your run: attach the JSON to your feedback email and I can load your exact state, see what you saw, and debug or improve from the real thing instead of a screenshot. The run file can contain everything you entered and everything the chain generated, so review it before sending and remove anything you wouldn't otherwise share. Notes alone, without a run file, are a complete response too.

The feedback that actually helps

You don't need a structured report. A few honest sentences on any of these is a full contribution:

Done? Send it in.

Email your notes, with the reviewed run file attached if you're willing, with subject "PDLC beta". I read and reply to everything, and a single pass with a few honest reactions is exactly what I asked for.

rob.dull@gmail.com

This is a working proof of concept suitable for portfolio review and limited evaluation with fictional, public, or sanitized information. It is not suitable for confidential enterprise work or production use.