Chapter 7 – Calibration: whose numbers, and why trust them
The Quantitative Schedule Risk Analysis Handbook · Edition 1, August 2026 · by the team behind [Epoch SRA](https://epochsra.com), an Excel add-in for schedule risk analysis · corrections welcome at contact@epochsra.com
A Monte Carlo simulation is arithmetic on assumptions. The arithmetic is never wrong; the assumptions usually are. Calibration is the discipline of forcing those assumptions to answer to reality – and its absence is the quiet scandal of most schedule risk practice.
The standard input, and its problem
Most SRA guidance says: ask the engineer for minimum, most-likely, and maximum durations, fit a triangle, simulate. The output looks rigorous. The input is the same optimism Chapter 1 described, now wearing error bars. Decades of estimation research say expert ranges are too narrow – actuals fall outside stated min-max far more often than the stated confidence implies. Simulating uncalibrated ranges produces distributions that are precise, plausible, and systematically tight: the P80 you compute is the world's P60, and you find out at the worst possible time.
What calibration means, concretely
Take historical programs where the outcomes are known. Replay them through the model: with the priors you propose, would the claimed percentile bands have contained the actual outcomes at the claimed rate? If your P20–P80 band is honest, roughly 70% of actuals should land inside it – not 95% (bands too wide, uselessly conservative), not 40% (too narrow, dangerously confident). Tune the priors until the claim and the history agree. That closed loop – claim, replay, tune – is the entire idea. A model that has never been replayed against outcomes has a precision claim and no accuracy claim.
The objections, taken seriously
"Calibrated against whose programs?" A fair question with an honest answer: whatever history was available – which is always a sample, always from somewhere. The defense is not that the sample is perfect; it is that the alternative input is calibrated against nothing at all. A rough map, honestly labeled, beats a confident blank page.
"Every program is unique." True at the task level, mostly false at the distribution level. Programs are unique in content and depressingly similar in failure statistics – optimism, [merge bias](/guide.html#merge-bias), supplier slips, test-retest cycles. This regularity is why reference-class forecasting works anywhere. And the objection contains its own endgame: as your organization accumulates its own actuals, the same replay loop can run on your history, and uniqueness flips from objection to asset.
Provenance: the label that earns the trust
Calibration evidence is never uniform across a distribution. The middle percentiles, where most outcomes land, can be disciplined by modest history. The deep tails – P90, P95 – rest on distribution-shape assumptions that thin samples cannot check. An honest model says so, per number: this percentile is backed by replayed outcomes; that one is a model estimate. In front of a review board, "here is which numbers carry evidence" beats "trust the software" every time it is tried – because it converts the board's skepticism from an attack into the model's own stated structure.
--- In practice: find one finished project in your organization with a preserved baseline. Compare its final actual finish to the range anyone would have stated at baseline time. That single data point, honestly faced, is the beginning of a calibration set – and usually the end of confidence in bare expert ranges.