The Delivery Contract
Moving an organization

Training teaches people.
Rollout changes an org.

Most AI adoption programmes are training programmes wearing a rollout's clothes. They reach a room, the room goes home, and nothing structural changed. The difference is whether the mechanism produces its own capacity — and whether anyone measured the people it was actually built for.

The reporting gap

Everyone can tell you the seat count

Most organisations adopting AI-assisted development can say how many seats they bought and how many people logged in. Very few can say what changed about the cost of shipping software. That is not a reporting problem. It is a design problem.

What the dashboard reports

  • Seats purchased and seats activated
  • Suggestion acceptance rate
  • Token or request spend
  • Sessions per developer per week

Every one of these can climb while the cost of getting a feature reviewed, merged, and released sits exactly where it started.

What the business bought

  • Cost per merged unit of work
  • Cycle time from intake to production
  • Review burden per change
  • Whether the number moved, on your own baseline

One question, asked in a form a finance function recognises.

Completed work is the only thing a business actually buys. Everything else on the dashboard is an input wearing a result's costume.

The unit of rollout

Pod, zone, wave

Scale does not come from bigger rooms. A room of two hundred watching a demonstration reaches nobody. Scale comes from a repeatable unit and a mechanism that manufactures more of the thing that constrains it.

Pod

Five people, one real repository, one real backlog item, merged by the end of the session. Not a sandbox. Not an exercise.

Zone

A set of pods with one dedicated floating facilitator. The facilitator is the constraint on everything — never the room, never the material.

Wave

One scheduled slot across all concurrent zones. Waves repeat; each one is expected to leave behind more facilitators than it consumed.

Delivery modePods per facilitatorDevelopers per zoneWhy
Remote420A stuck pod is discovered at a checkpoint, minutes after it stalled
Onsite630A stuck pod is visible at a glance, in seconds

That difference — reading a room at a glance versus discovering a stall at a checkpoint — is the entire argument for travelling. Onsite is an accelerant, not a prerequisite. If it becomes a prerequisite, the programme cannot scale past your travel budget.

Self-propagation

The most consequential decision is how many zones you start with

Every pod is expected to produce two qualified facilitators for the next wave, with manager-confirmed availability. That single expectation is what turns a training programme into a rollout mechanism — and it means starting capacity compounds instead of merely repeating.

Coverage projection

Interactive
Session 3 mode
4waves at 3 zones
6.5months to coverage
6waves at 1 zone
+3.3months lost to starting small
Starting with 3 zones Starting with 1 zone
Cumulative developers reached by wave, comparing a multi-zone start against a single-zone start.
Cumulative developers reached, by wave. Assumes seven weeks between waves, two qualified facilitators produced per pod, and roughly a third converting to manager-confirmed availability. Coverage is capped at the stated population.
View as table

Attendance does not count as capacity

A graduate is someone whose manager has confirmed they will facilitate a future wave. Not someone who showed up and enjoyed it. Programmes that skip this line report healthy numbers for three waves and then quietly stall, because the operator bench was always notional.

Under one zone, you never stop being the constraint

With a single starting zone the programme stays facilitation-bound from the first wave to the last — you are the bottleneck for its entire life. Start with three and trained internal capacity overtakes the remaining population partway through, at which point the organisation is running its own waves and you are optional. That is the goal.

Designing for the spread

The room contains both ends of your organisation

Someone who has shipped agent-driven work will sit next to someone who has used AI as autocomplete and nothing else. That spread fails in both directions at once: the top disengages because it is too slow, the bottom goes silent because it is too fast.

The question is not how to teach the bottom 80%. It is how to run a room where the top 20% are present and the bottom 80% still does the work.

Seed each pod with a strong operator — hands off the keyboard

Their role is defined out loud as facilitator, not driver. They do not touch the keyboard for the duration. This converts your biggest disengagement risk into distributed facilitation capacity, and it is what lets a room run more zones than you have staff for. The strongest engineers find this harder than the exercise, which is itself the lesson.

Name the driver; rotate on a timer

The keyboard moves at fixed intervals and the current driver is named. Nobody watches for four hours and nobody quietly opts out. Small mechanism, most of the work.

Gate on prerequisites

Prerequisites do not eliminate the spread — they raise the floor and compress it, which is enough. A pod where everyone has run an agent once successfully starts on the problem instead of on setup.

Prepare a stretch track

Multi-agent orchestration, custom agent definitions, evaluation harnesses, policy as code. A pod that finishes early gets the next thing, not a coffee break. The strongest people should leave stretched even though they were never the target.

Measurement

Three tiers, and the average is not one of them

An average can improve while the group you built the programme for does not move at all. That is exactly what happens when the top quintile gets better and everyone else stays put — and it is the most common way an adoption programme reports success it did not have.

Tier 1 — Completed work

  • Cost per merged unit of work
  • Cycle time, intake to production
  • Review burden per change

The numbers leadership actually asked for.

Tier 2 — Adoption reality

  • Developers at zero activity
  • Split into no access, hit a limit, chose not to
  • Silent stalls

Consistently surfaces problems the utilisation dashboard reports as success.

Tier 3 — Governance health

  • Share of agentic changes passing review unmodified
  • Where guardrails are being hit
  • Bypass rate

A pipeline nobody trusts gets bypassed — and bypass is measurable.

Silent stalls

An engineer hits a constraint — a limit, a permission, a broken environment — and quietly stops rather than asking for more. They stay counted as an active seat. They stop being a participant. This is the single most under-measured failure in AI adoption, and it is why "developers at zero" has to be split by cause, not just counted.

Report the distribution, not the mean

Developers at zero

Same cohort, before and after. The number that matters most.

Bottom-two-quartile movement

The shift in the group your existing reporting already identifies.

Distribution shape

If the spread widened, you reached the people who were already fine.

The question-quality signal

"Are we getting the questions we should be getting?" sounds soft and is not. Categorise questions asked in-session; the ratio shift between the first session and the last is the cleanest learning signal available, and it costs nothing to capture.

BucketWhat it sounds likeWhat it means
Tool mechanicsHow do I make it do the thingStill operating the tool
WorkflowHow should we structure this workStarting to direct it
Architecture & governanceWhat should we let it decide, and where do we draw the lineThe shift has landed

A room that starts in the first bucket and ends in the third has moved. The downstream version of the same signal: count inbound architecture and design-review requests before and after. When teams start bringing design questions instead of code questions, the change is real.

Measurement has to be tool-neutral

In any organisation running more than one assistant, seat-level and token-level comparisons mislead by construction — unlimited plans hide cost, metered plans expose it, and neither figure says what was produced. Attributing cost to completed work regardless of what produced the diff is the only fair comparison, and it happens to answer the question leadership actually has.

Runway

Lead time fills the room; compression empties it

The date is chosen to fit the runway, not the runway compressed to fit the date. This is the least glamorous section here and the one that most reliably decides whether the working session works.

3 months

Full room. No excuses left standing.

2 months

Most people. Some conflicts you cannot recover.

30 days

About half. And the half you lose skews toward the people you were trying to reach.

MilestoneTiming
Internal review with engineering leadersT − 10 weeks
Date published; invite and prerequisite packet issuedT − 9 weeks
Session 1 — the shift, 90 minT − 4 weeks
Session 2 — a pipeline, inspected, 90 minT − 2 weeks
Prerequisite completion checkpointT − 2 weeks
Repositories pre-flighted, backlogs refined, machines verifiedT − 1 week
Session 3 — pods build and mergeT
Review 1 — what shipped, what stalledT + 3 weeks
Review 2 — measured before and afterT + 7 weeks

The practical test for any candidate date: count back nine weeks and ask whether the internal review and the published invite can realistically land before that point. If not, moving the date is cheaper than absorbing a half-full room.

One decision to make deliberately, not by default

Are prerequisites advisory or a condition of attendance? Advisory is easier to communicate. A condition produces a materially better room. Either is workable — but it has to be chosen and stated in the invite, because defaulting to advisory silently is how a working session becomes a demonstration.

Failure modes

Named plainly, because they are all real

Every one of these has cost someone a working session. None of them are exotic; all of them are cheap to prevent and impossible to fix on the morning.

The rescue spiral

The pod arrives without prerequisites, stalls, and the experienced operator picks up the keyboard to save it. The session quietly becomes a demonstration. This is the failure mode — the checkpoint two weeks out exists to catch it while there is still time to chase people.

No merge authority in the room

Pods build something real and cannot merge it. The single deliverable of the day does not exist. Confirm it per pod, by name, before the date is published.

Theatre seating

Pods need to sit around something. Not negotiable — it is what determines whether the day is a workshop or a talk. Ask what the tables look like, not what the room capacity is.

Untested network

Guest wifi that blocks your platform or your model endpoints ends the session. Test from the actual room on the actual network before the day. Assumed is not tested.

The first hour lost to setup

Machines provisioned, access confirmed, one successful agent run — in advance. An hour is a quarter of the day, and it is always the same hour.

Nobody owns the before-and-after

One named person runs the cohort snapshot before Session 1 and again at Review 2. Without that name, the programme gets assessed on impressions — and impressions always favour the top 20%.

Better to run the day with sixty people who did the preparation than a hundred who did not.

What the organisation is left with

Independent of you, at the end

A governed pipeline

Running in their own repositories, with guardrails their teams chose — built to be copied, not admired.

An operator bench

Large enough to run further waves without you. This is what makes a recurring cadence realistic.

A measured delta

Cost per completed feature, before and after, on their own baselines and their own cohort report.

A rollout mechanism

Headcount growth becomes a scheduling question instead of a new programme.

The outcome is not that developers use AI more. It is that the organisation can say, with evidence, what it now costs to ship a feature — and whether that number moved.