Delivery Forecast
When will it be done? This is a Monte Carlo delivery forecaster: it answers with a range and the odds attached to it, instead of a single date.
How it works. You give it how much work your team actually finished in each of the last several periods — weeks, sprints, days, whatever you count in — and how much is left. It then plays your project forward 10000 times. Each run draws periods at random from your own history, one after another, until the remaining work is used up, and records how long that took. Some runs draw a string of good periods and finish early; some draw the bad ones and drag. The chart below the form is all 10000 of those finishing times at once.
Why a range replaces a date. A single ship date is a forecast with the uncertainty deleted — it is one draw from that chart, presented as though it were the whole chart. It is usually close to the middle, which means it is roughly a coin flip, and it carries no way to say how wrong it might be. A distribution keeps what the date threw away: P50 is the coin flip, P80 is the answer you can commit to, and the gap between them is your real risk, in periods, on the page.
There is no model in the loop. Nothing on this page calls an AI system, and none of the numbers you are about to see came from one. It is arithmetic and a seeded random number generator — the same Go code the tests cover, compiled to WebAssembly. That is deliberate: a forecast you cannot reproduce is a forecast you cannot argue with, and every run here is reproducible from its inputs and its seed. A language model would be genuinely useful one step earlier, turning a messy issue-tracker export into the throughput numbers this form wants; it would not make the forecast itself better, and we would rather say that than imply otherwise.
Everything runs in this tab. Your throughput history is typed into this page, simulated by Go compiled to WebAssembly, and never transmitted — there is no endpoint behind this form, and no request is made after the page and its module have loaded. Disconnect the network once it is ready and it keeps working.
Loading the simulation module…
- P50 — half of the simulated futures are later than this
- P80 — one in five is later
- P90 — one in ten is later
Horizontal axis: periods. Bar height: how many of the simulated futures finished in that range.
What this cannot know
Sampling your past periods at random assumes next period is drawn from the same distribution the last ones were. A team that just lost half its people, or just finished the hard part, breaks that assumption, and nothing in the arithmetic can detect it. The output is a statement about the history you typed, not about your project.
It also does not know what you have not thought of yet. The scope-growth control is a crude stand-in for discovered work, and the number in it is yours to choose — we have not measured your scope growth and cannot.
The module is 3.3 MB uncompressed. A Go binary carries its garbage collector and scheduler whether a program uses them or not, so a hand-written JavaScript implementation would be a fraction of the size. It buys the guarantee that the simulation on this page and the simulation the tests cover are the same code.
We graded it against our own delivery history, and it missed
A forecaster nobody has checked is a random number generator with a chart. So this one was run against the only delivery record this site owns: 270 development sessions logged over 36 days, 25 July 2026 to 29 August 2026. The grading is walk-forward — at each point the forecast sees only the days before it — and asks how long until 10 more sessions closed. That gives 27 gradeable outcomes; 2 more are excluded because the record ends before they resolve.
| Band | Claims to contain | Actually contained | Outcomes |
|---|---|---|---|
| P50 | 50% | 48.1% missed | 13 of 27 |
| P80 | 80% | 63.0% missed | 17 of 27 |
| P90 | 90% | 63.0% missed | 17 of 27 |
The arithmetic is not what failed. Run against data that genuinely is drawn from one distribution, the same code is conservative — its bands contain more than they promise, never fewer. What broke here is the assumption, and the break is visible in the data: throughput ran at 10.64 sessions a day for the first 22 days and 2.57 a day for the last 14, and never recovered. A method that draws past periods at random cannot represent “and then it stopped being like that”.
That is what the history-window control is for, and here is exactly what it buys. Narrowing the window improves every band. Every width still misses. And it is not a case of narrower being better — past a point, a shorter window is a smaller sample, and the variance it adds costs more than the staleness it removes.
| Window | P50 | P80 | P90 | Outcomes outside a band |
|---|---|---|---|---|
| all history | 48.1% | 63.0% | 63.0% | 34 |
| last 14 periods | 48.1% | 66.7% | 74.1% | 30 |
| last 10 periods | 51.9% | 74.1% | 77.8% | 26 |
| last 7 periods | 66.7% | 74.1% | 81.5% | 21 |
| last 5 periods | 66.7% | 70.4% | 81.5% | 22 |
Every miss, uncurated
All 34 of them, sampling the whole history at a 10-session horizon. Nothing is filtered and nothing is summarized away — a calibration figure with its misses hidden is a marketing number.
| Band | Forecast from | Days of history | Said | Took |
|---|---|---|---|---|
| P50 | 2026-08-01 | 7 | 1 periods | 2 periods |
| P50 | 2026-08-03 | 9 | 1 periods | 3 periods |
| P80 | 2026-08-03 | 9 | 2 periods | 3 periods |
| P90 | 2026-08-03 | 9 | 2 periods | 3 periods |
| P50 | 2026-08-04 | 10 | 2 periods | 3 periods |
| P80 | 2026-08-04 | 10 | 2 periods | 3 periods |
| P90 | 2026-08-04 | 10 | 2 periods | 3 periods |
| P50 | 2026-08-16 | 22 | 2 periods | 6 periods |
| P80 | 2026-08-16 | 22 | 2 periods | 6 periods |
| P90 | 2026-08-16 | 22 | 3 periods | 6 periods |
| P50 | 2026-08-17 | 23 | 2 periods | 6 periods |
| P80 | 2026-08-17 | 23 | 2 periods | 6 periods |
| P90 | 2026-08-17 | 23 | 3 periods | 6 periods |
| P50 | 2026-08-18 | 24 | 2 periods | 6 periods |
| P80 | 2026-08-18 | 24 | 2 periods | 6 periods |
| P90 | 2026-08-18 | 24 | 3 periods | 6 periods |
| P50 | 2026-08-19 | 25 | 2 periods | 5 periods |
| P80 | 2026-08-19 | 25 | 2 periods | 5 periods |
| P90 | 2026-08-19 | 25 | 3 periods | 5 periods |
| P50 | 2026-08-20 | 26 | 2 periods | 4 periods |
| P80 | 2026-08-20 | 26 | 3 periods | 4 periods |
| P90 | 2026-08-20 | 26 | 3 periods | 4 periods |
| P50 | 2026-08-21 | 27 | 2 periods | 3 periods |
| P50 | 2026-08-22 | 28 | 2 periods | 3 periods |
| P50 | 2026-08-24 | 30 | 2 periods | 5 periods |
| P80 | 2026-08-24 | 30 | 3 periods | 5 periods |
| P90 | 2026-08-24 | 30 | 3 periods | 5 periods |
| P50 | 2026-08-25 | 31 | 2 periods | 5 periods |
| P80 | 2026-08-25 | 31 | 3 periods | 5 periods |
| P90 | 2026-08-25 | 31 | 3 periods | 5 periods |
| P50 | 2026-08-26 | 32 | 2 periods | 4 periods |
| P80 | 2026-08-26 | 32 | 3 periods | 4 periods |
| P90 | 2026-08-26 | 32 | 3 periods | 4 periods |
| P50 | 2026-08-27 | 33 | 2 periods | 3 periods |
Is this a fair test?
Partly, and where it is not, it is not. It is a real test of the arithmetic against a real, bursty arrival process that nobody generated to be forecast. It is not a software backlog: a development session is a unit of attention rather than of scope, so “sessions remaining” was never a quantity anybody was burning down. The completion time is also a proxy — the last log line carrying a session's tag — which runs early for a session that ended without writing one. And 27 outcomes is a small sample. We are showing it because it is the delivery record we actually have and grading against it is more honest than grading against nothing.
The grading code, the raw session list and the numbers in these tables are all in the repository, and the tables above are computed by running that code at render time rather than transcribed from a test log.