19. Training, Cadence, and Failure Recovery

On a Friday six weeks in, the rollout still looked healthy.

Then one lead change, one urgent supplier issue, and one skipped review cycle later, the old behavior was back — a program quarter of uncontrolled drift before the cadence recovered. The relapse analysis found three gaps: the skipped review had been covered by a handoff email, not a backup-trained lead; the lead change had no ownership-transfer protocol; the supplier issue had triggered an exception meeting that displaced the review cadence without a recovery date.

Three controls were added: a named backup DRI for each cadence slot, a thirty-day cadence audit after any lead change, and an explicit recovery trigger — if any scheduled review is displaced, the next one is held within five days regardless of program pressure. Six months later, a second lead change hit the same team. The cadence held.

Training is behavior transfer, not slide transfer

New people do not need the full theory first.

They need operational moves in context.

Onboarding minimum:

  1. ownership map and closure expectations,
  2. one-page truth mechanics and source links,
  3. risk trigger/escalation behavior,
  4. decision record standard.

Teach these on live program threads, not in abstract slides only.

Cadence is the enforcement mechanism

Without cadence, standards drift into preference.

The minimum cadence is four touchpoints. Any program that cannot hold all four is not running the OS — it is running a lighter version that will drift when pressure spikes:

  • weekly decision/risk sync (updated decision log and risk register entries),
  • fixed one-page refresh rhythm (published one-page status before each review),
  • gate prep/check with artifact traceability (gate artifact delta from the record system),
  • monthly retro on control-loop failures (playbook update with named control adjustments).

Cadence should be predictable and short enough to survive workload spikes.

Cadence is the mechanism that makes the artifact trustworthy enough to read between meetings. The engineering lead who has locked a working cadence has bought the program a habit that survives the next bad week.

Failure recovery without blame loops

When the system slips:

  1. identify the control that failed (owner, record, gate, risk, status),
  2. identify why it failed (capacity, ambiguity, missing authority, poor handoff),
  3. restore minimum behavior on one active thread,
  4. capture adjustment in playbook.

Do not frame recovery as "who failed process."

Frame it as "which control failed under what condition."

Metrics for real adoption

Activity metrics mislead when closure and mismatch outcomes stay flat.

The rollout chapter defines each metric and its artifact source. This chapter's job is reading the trend over consecutive cycles — a single-cycle improvement can be noise; a sustained two-cycle trend is the signal that behavior is actually changing, not just reported:

  • decision closure cycle time — trending down, or plateaued after initial improvement?
  • reopen rate on closed decisions — still dropping, or holding flat while closure rate rises?
  • one-page vs source mismatch frequency — gone, or recurring on the same artifact types?
  • late risk discovery rate — at zero after two cycles, or spiking around milestones?
  • unresolved ownerless decision count — cleared each cycle, or accumulating?

If these improve for two reporting cycles, adoption is real.

If only template completion improves while outcome metrics stay flat, the rollout produced activity, not behavior change.

Turnover resilience

Systems fail at role transitions unless handover is explicit.

Require handover package for key roles:

  • active decisions and owners,
  • top risks and triggers,
  • current one-page state,
  • unresolved escalations,
  • next gate commitments.

This turns personnel changes into manageable events.

Keep the bar practical

Overly rigid process breaks under load spikes.

Overly loose process breaks when ownership or data is ambiguous.

Durable cadence means:

  • strict on ownership and traceability,
  • flexible on meeting format and local workflow details.

Eight weeks later, same program. The lead changed. The supplier issue hit. The review cycle slipped by one day, not three weeks — because the one-page had been published before Friday, the risk register named an owner on the supplier item, and the incoming lead read both before the first review. The handover package was on the shared drive: active decisions, top risks, current one-page, unresolved escalations, next gate commitments — the incoming lead was fully current within one review cycle.

Field test: name the last time a control loop on your program failed — a gate arrived without evidence, a DRI was absent, a cadence slipped. Is there a playbook entry that names the trigger, the recovery action, and the owner? If not, recovery is improvised every time, and the same failure will return.


The control stack is not a substitute for regulated lifecycle frameworks where the audit trail, evidence standard, or approval chain is defined by statute rather than by the program.