12. Models and Tests That Change Decisions
Wednesday, 6:40 p.m. The decision review packet holds 20 model runs and 12 test reports and still cannot resolve the thermal case temperature requirement.
The gate decision still hangs because no owner can show baseline, delta, and effect size in one traceable artifact.
Evidence volume is high, but decision power stays low when each run is not linked to a baseline requirement state. The missing discipline is not more runs — it is calibrating the model against the chamber and expressing the result as a delta from baseline: what changed, by how much, and what that means for the open requirement. That calibration-and-delta loop is what this chapter installs.
Evidence must be delta-first
Your model or test result is useful only relative to a named baseline artifact — without the baseline, you cannot tell whether the new run changed anything or just confirmed that you ran it again.
For every run record, capture a decision delta package:
- cite baseline reference,
- state what changed,
- predict expected effect,
- report observed effect,
- declare decision implication.
Without this package, results become isolated artifacts and cannot form a learning sequence that changes decisions.
Naming and traceability discipline
Use one boring but enforceable traceability convention:
- keep stable requirement ID linkage,
- assign run and test IDs,
- record revision IDs for design and setup,
- stamp timestamp and owner,
- store records in a location with immutable history.
This discipline prevents decision debt: review leads pay the cost at gate when they cannot compare outcomes across builds and teams.
Pair model claims with physical evidence
Models are strongest when analysts calibrate them against physical evidence. Tests are strongest when reviewers interpret them with model context and assumptions.
Use these pairing questions before changing a requirement value:
- Where do model and test agree?
- Where do they diverge?
- Which assumptions explain divergence?
- What decision can be made now despite remaining uncertainty?
The goal is not perfect correlation before every move; the goal is honest confidence for the next dated decision.
A thermal model for a power-electronics module predicted the switching-stage case temperature would hold at 73 °C at full duty — safely below the 75 °C thermal case temperature limit. The chamber test measured 78 °C steady-state. The team's first interpretation was that the chamber was measuring something different than the model. The calibration finding was narrower: the model's contact-resistance boundary condition assumed a thermal-grease application that the actual mounting jig could not deliver. The jig had been designed for a different module with different interface geometry — it had been on the test floor for three years, and nobody had updated the boundary-condition assumption when the module changed. The model was correct given its boundary condition. The boundary condition was wrong given the actual assembly. Treating the model and the chamber as independent confidence votes is what stretched this out: a full week of "the chamber must be wrong" before anyone examined the boundary condition. One boundary-condition correction and a re-run later, the model predicted 79 °C — matching the chamber within measurement uncertainty, a delta of one degree where there had been five. What changed was not the model and not the chamber but one entry in the boundary-condition log: the contact-resistance value, now traced to a direct measurement on this module instead of carried over from a different one. The requirement revised to 78 °C — a controlled update: the correction is the evidence, the new value is documented with a named owner, and the baseline moves.
That revision is a controlled evidence update — the requirement baseline moves only when evidence satisfies predeclared criteria and the rationale is documented.
Evidence threshold for requirement changes
Do not revise a requirement on a single model run that shows a passing result.
Set requirement-change thresholds in advance:
- confirm the test or model ran at the operating conditions the requirement specifies,
- verify repeatability is acceptable,
- show sensitivity to key variables is understood,
- bound measurement uncertainty enough for the decision at hand.
If threshold is not met, the evidence owner documents what is missing in the review artifact and sets a closure date.
Common evidence failure patterns
- orphan files with no requirement linkage,
- results impossible to reproduce from metadata,
- "best run" selection without rationale,
- model updates with no calibration note,
- test deltas reported without setup changes.
If you can tick two or more of these in your current evidence folder, the gate review will be a debate, not a decision.
Each pattern inflates confidence in the artifact while leaving decision uncertainty unresolved at review.
Practical evidence package for decisions
For any decision review, provide one compact decision package:
- state question to answer,
- summarize baseline and delta,
- rate evidence quality,
- recommend a decision,
- assign residual risk and next evidence step.
This package keeps technical depth available while giving exec and PM reviewers a clear, fair decision frame.
The boundary-condition correction is confirmed. The model re-run matches the chamber within measurement uncertainty. The gate decision changes from hold to proceed — one week after the chamber result that everyone initially attributed to measurement error. The requirement is 78 °C, the model predicts 79 °C, and the boundary-condition log now carries the contact-resistance value traced to the direct contact-resistance measurement. The decision packet is one page. The technical backup is three.
Field test: name one model your program relies on. Did the last model update change a decision, or did it confirm a decision already made? If it confirmed what was already decided, the model is not driving decisions — it is defending them. Find the last result that was acted on and trace it to the decision log.
The one-page packet settled the requirement for one unit, on one bench, on one Wednesday. It said nothing about the forty units already coming off the line behind it.