How wrong we are, and where.
The LiftLine AI forecast scoreboard: error by predicted crowd level, per lift, in minutes and in people, with the coverage denominator and the known weaknesses stated.
This is the scoreboard. It covers one mountain, one quiet spring, and a small fraction of the forecasts we issued. Read the first table before the headline number — the headline is carried almost entirely by the hours when nothing was happening.
When we said it would be busy, we were wrong
Forecasts are grouped by what the model itself said the crowd level would be, then scored against what the cameras and skier reports saw. This is the table we would want to read about someone else's forecast, so it goes first.
| We said | Hours | Average error | Median error | p90 error | Bias | The line actually ran | We said it would run |
|---|---|---|---|---|---|---|---|
| Quiet | 425 | 1.85 min | 1 min | 5 min | +1.56 min | 1.94 min | 3.51 min |
| Moderate | 95 | 9.85 min | 10 min | 15 min | +9.85 min | 2.73 min | 12.58 min |
| Busy | 34 | 17.47 min | 17.5 min | 25.7 min | +17.47 min | 2.41 min | 19.88 min |
When we say a line is quiet we are usually right. When we say it is busy we are usually wrong, and wrong by about eighteen minutes — on exactly the hours a rider would change their plan over.
The pooled number, in context
Scored against camera-derived waits at Mammoth, the day-of forecast has a mean absolute error of 4.18 minutes across 554 forecast-hours over 100 lift-days at Mammoth, April 5 to June 10, 2026. Over those same hours the average line was 2.11 minutes long, so most of that error is us over-calling a quiet lift. It covers 12.7 percent of the forecasts issued in that window; the rest were never observed and so were never scored.
The window, and the number without it
A lift-hour is scored only when (a) every camera observation behind it was produced by the wait conversion currently in production, and (b) it falls inside the period of continuous scheduled collection.
The published figure has to answer one question: how well does the system that exists today forecast? Observations produced by a conversion that has since been deleted cannot answer it, and neither can days on which the collector was not running.
| Lift-hours | Lift-days | Average error | Bias | The line actually ran | |
|---|---|---|---|---|---|
| Scored window, 2026-04-05 to 2026-06-10 | 554 | 100 | 4.18 min | +3.96 min | 2.11 min |
| Every scored hour on record, 2026-03-19 to 2026-08-23 | 679 | 129 | 5.34 min | +5 min | 2.44 min |
The whole scored record with no window applied, published so the windowed figure can be checked against what it was cut from. This is NOT the headline. It mixes three ground-truth conversions and counts 2026-05-08 three times.
| Month | Lift-hours | Average error | Bias | In the window |
|---|---|---|---|---|
| 2026-03 | 29 | 20.69 min | +20.62 min | no |
| 2026-04 | 546 | 4.15 min | +3.74 min | partly — 2026-04-01 to 04-04 excluded |
| 2026-05 | 44 | 7.25 min | +7.21 min | yes |
| 2026-06 | 3 | 1 min | +1 min | yes |
| 2026-08 | 57 | 7.68 min | +7.61 min | no |
The window removes 125 scored hours, taking the record from 679 hours to 554. This is every hour it removes, and the rule that removes it.
- 2026-03-19 to 2026-03-31
- 29 scored hours. Ground truth produced by the pre-Q/mu conversion; only two of the five scored cameras existed; no scheduler. This is the build period. This block is the single largest mover of the headline. It is removed on the conversion-change rule, which was fixed before the per-month figures were looked at.
- 2026-04-01 to 2026-04-04
- 39 scored hours. Ground truth still produced by the M/D/1 formula (deleted 2026-04-04 12:11) or by the pre-model path, and the camera set was still changing. Removing these four days makes the headline WORSE than a plain 'drop March' cut would (4.18 against 4.25). They are removed because the rule says so.
- after 2026-06-10
- 57 scored hours. Outside continuous collection — the scheduler had been off for two months. These are off-season dev runs.
| Start date | What supports it | Lift-hours | Average error | Bias |
|---|---|---|---|---|
| 2026-04-01 | calendar April, no event behind it | 565 | 4.25 min | +4 min |
| 2026-04-05 | CHOSEN — first whole day on the current conversion with all scored cameras live | 554 | 4.18 min | +3.96 min |
| 2026-04-06 | rejected — no event supports it; it merely drops a bad day (2026-04-05 scored 14.61) | 521 | 3.52 min | +3.29 min |
Published so the boundary can be second-guessed. The chosen start is 2026-04-05 because that is where the evidence lands. It is NOT the start date that produces the best number — starting one day later would improve the headline by 0.66 minutes, and that day was kept.
Quoting 4.18 minutes on its own would be technically true and substantively misleading. Against a line that averaged 2.11 minutes, an error of 4.18 minutes is not precision — it is us calling a crowd that was not there.
Both series, including the one that is not a forecast
| Series | What it is | Horizon | Lift-hours | Lift-days | Average error | Median | Bias | p90 |
|---|---|---|---|---|---|---|---|---|
| Day-of AI forecast | The AI forecast, written each morning and never touched again | Same day, issued around 01:11 local | 554 | 100 | 4.18 min | 2 min | +3.96 min | 12 min |
| Live-corrected estimate — not a forecast | The current-hour number after it has already read the cameras | None | 84 | 17 | 6.19 min | 5 min | +6.12 min | 14.7 min |
| Mountain sim, day-ahead | Not scored yet | Day ahead | 0 | 0 | — | — | — | — |
| Sim after live update | Not built yet | None | 0 | 0 | — | — | — | — |
How little of this gets scored
A forecast-hour can be scored only where a camera or a skier saw that lift in that hour. In the published window we issued 4,346 lift-hours of forecast and could check 554 of them — 12.7 percent. Across the whole record 6 lift lines at one resort carry a camera; 5 of them recorded frames inside this window, and 9 lifts were scored — because a lift can be scored from skier reports with no camera on it at all. Every other lift in the app is a forecast nobody has ever checked.
| Issued | Scoreable | Coverage | |
|---|---|---|---|
| Forecast-hours | 4,346 | 554 | 12.7% |
| Lift-days | 486 | 100 | 20.6% |
| Lifts | 16 | 9 | — |
More lifts were scored (9) than carry a camera (5) because four more were scored on a handful of hours from skier reports alone.
Per lift
Only five lifts have enough scored hours to say anything about.
| Lift | Hours | Days | Average error | Median | p90 | Bias | Line ran | We said |
|---|---|---|---|---|---|---|---|---|
| Broadway Express | 149 | 26 | 5.62 min | 3 min | 14 min | +5.51 min | 1.92 min | 7.43 min |
| Discovery Chair | 147 | 23 | 3.39 min | 3 min | 7 min | +2.84 min | 4.56 min | 7.4 min |
| Canyon Express | 97 | 17 | 5.44 min | 1 min | 16.8 min | +5.38 min | 0.91 min | 6.29 min |
| Village Gondola | 87 | 16 | 3.62 min | 1 min | 12 min | +3.62 min | 0 min | 3.62 min |
| Face Lift Express | 68 | 12 | 1.9 min | 1 min | 5 min | +1.66 min | 1.68 min | 3.34 min |
Four more lifts — Gold Rush Express, Stump Alley Express, Roller Coaster Express, Chair 23 — have between 1 and 2 scored hours each. That is not a sample, so we are not printing a number next to them.
Is the forecast wrong, or is the thing measuring it wrong?
Ground truth is not measured. Cameras count heads; a queue model divides by an asserted service rate to get minutes. If that service rate is too generous, the derived 'actual' is too low, and a correct forecast would look like constant over-prediction. The observed bias is +3.96 and remarkably one-directional, which is the signature a biased ruler would leave.
| Service rate too generous by | Average error | Bias | The line would have run |
|---|---|---|---|
| 1x | 3.567 min | +3.347 min | 2.073 min |
| 1.5x | 3.414 min | +2.304 min | 3.116 min |
| 2x | 3.739 min | +1.218 min | 4.202 min |
| 2.6x | 4.447 min | +0.004 min | 5.416 min |
| 3x | 4.98 min | -0.855 min | 6.275 min |
The service rate that best FITS the data is about 1.5x too generous, which is consistent with the Bayesian learner having already cut it 29-47% where it had enough drain events to learn. That correction removes about 31% of the bias. Erasing the bias entirely needs 2.6x, which is a different claim and a much larger one.
| Lift | Nameplate | Rate we use | Rate needed to clear us |
|---|---|---|---|
| Broadway Express | 3,000/hr | 905/hr (30.2%) | 348/hr (11.6%) |
| Discovery Chair | 2,400/hr | 448/hr (18.7%) | 172/hr (7.2%) |
| Canyon Express | 3,000/hr | 1,055/hr (35.2%) | 406/hr (13.5%) |
| Village Gondola | 3,600/hr | 1,924/hr (53.4%) | 740/hr (20.6%) |
| Face Lift Express | 2,400/hr | 971/hr (40.5%) | 373/hr (15.6%) |
The rates already in use imply the lifts run at 19-53% of nameplate while a line exists, which is low but arguable. The rates needed to erase the bias imply 7-21% — a detachable six-pack loading 348 people an hour with a queue in front of it. That is not a service rate, it is a broken lift.
Hours the ruler cannot be blamed for
Isolate in-window hours in which EVERY camera frame saw 10 or fewer people in the queue. On those hours the derived wait is 0 because of the zero-threshold rule, not because of a division — no service-rate error can change them. Compare the forecast's bias there against its bias everywhere else.
| Hours | Our bias | |
|---|---|---|
| Hours the service rate could explain | 448 | +4.14 min |
| Hours it cannot touch | 95 | +3.52 min |
On hours where the camera never saw more than about five people in line, the forecast still said 2.81 minutes, and it still says so after every correction the ruler could absorb. Twelve of those 91 hours were called moderate or busy. The over-prediction on hours the ruler cannot explain is 85% as large as on hours it could. That is not the signature of a biased ruler.
This is a bound, not a measurement. It rests on the counterfactual being computable from stored columns (it is: the conversion reproduces from live_crowd.queue_filtered and live_crowd.service_rate on 2,588 of 2,591 in-window frames, the three misses being 1-minute rounding). It cannot be closed without independently timed waits.
The same forecast, in the unit we actually measure
A camera counts heads. Minutes are inferred from heads by a queue model. So here is the same forecast scored in heads, using each lift-hour's own measured service rate.
| Value | |
|---|---|
| Lift-hours | 515 |
| What the camera counted | 23.87 people, on average |
| What the forecast implies | 74.36 people |
| Mean absolute error | 53.9 people |
| Median absolute error | 18.9 people |
| Bias | +50.49 people |
The forecast implies a line 3.12 times longer than the one in the picture. This is the wait error rescaled rather than independent evidence — but heads are the thing we measure and minutes are the thing we guess, so this is the honest way round.
What is wrong with our forecast
First person, no hedging, in the order that matters.
- The forecast runs high, almost always. Bias is +3.96 minutes against a mean observed wait of 2.11 minutes. 95 percent of the average error is one-directional over-prediction, not scatter.
- When the forecast says the line is busy, it is usually wrong. On hours the model labeled 'high', it predicted 19.9 minutes and the line ran 2.4. That is a 17.5 minute average error on exactly the hours a rider would act on.
- The number is built on a quiet spring. Every scored hour is from 5 April to 3 June at Mammoth. The mean line was 2.1 minutes long. Nothing here says anything about a Saturday in January.
- Almost nothing gets scored. 554 of 4,346 issued forecast-hours had any ground truth at all, which is 12.7 percent. Five lift lines at one resort have cameras. Every other lift in the app is unverified, including all 51 lifts forecast at the seven other resorts.
- The ground truth is itself a model output. A camera counts heads; a queue model turns heads into minutes using service rates that were asserted, not measured. The best-fitting correction to those rates is about 1.5x and accounts for roughly a third of the over-prediction. It does not account for the rest.
- Any line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29 percent of every camera frame inside the window.
- The live corrector barely corrects. On the 84 hours where both series exist it changed 12 of them and improved the average error by 0.42 minutes, and it did that by reading the same observations it is then scored against.
- Twelve and a half percent of the scored baseline hours come from prediction rows written after the hour they describe. Those are hindcasts sitting inside a forecast number.
- The published figure is a windowed figure. 125 scored hours were excluded by the window rule, including the 29 worst hours on record. The unwindowed number is 5.34 and is published beside it.
What the ground truth actually is
Our ground truth is 3,635 camera observations across 5 lift lines at mammoth (2026-04-05 to 2026-06-07), plus 13 skier-submitted waits across 7 lifts. It is not a stopwatch.
- The median frame saw 15 people in line; the largest saw 160.
- A line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29.2 percent of every frame we have.
- No stopwatch has ever been held against this model. Not once. The conversion from heads to minutes has never been checked against a timed wait.
Numbers we withdrew
These were published and are not true as stated. They are listed rather than deleted, because a correction that leaves no trace is not a correction.
| Withdrawn claim | Verdict | What replaced it |
|---|---|---|
| about 6 minutes of average forecast error across 138 scored lift-days | does not reproduce as stated | 5.99 minutes, the unweighted mean of the per-lift-day mae column across all 138 rows, mixing the leakage-free baseline series with the live_corrected series that consumed its own ground truth. |
| 5.12 minutes of mean absolute error, 622 forecast-hours, 19 March to 3 June | superseded by this document, not wrong | That figure was correctly computed for the window it named. It has been replaced because the window itself was calendar-shaped rather than event-shaped: it began at the first camera frame, which predates the wait conversion now in production by sixteen days. The comparable all-data figure recomputed here is 5.34 over 679 hours; the difference is 57 post-season hours the earlier pass did not include. |
| every forecast gets scored | false | 12.7 percent of issued forecast-hours were scoreable inside the published window. |
| 6,294 webcam crowd counts across 17 lift lines | does not reproduce | live_crowd holds 4,522 real camera rows across 6 Mammoth lift lines, plus 3,139 synthetic rows written by the simulator on 2026-08-17 and 2026-08-18. engine/vision/config.py defines exactly 6 LIFT_CONFIGS entries, all Mammoth. Neither 6,294 nor 17 reconciles. Flagged for the owner; this figure is outside this lane's scope to fix. |
Scope of everything above
- Resorts
- mammoth — Mammoth is the only resort with any ground truth. Every other resort in the catalog has predictions and zero observations.
- Window
- 2026-04-05 to 2026-06-10, 29 days with predictions
- Excluded
- resort_slug = 'simulator' — Synthetic data generated by the mountain sim on 2026-08-17 and 2026-08-18 under mctd_ids 9001-9011. Not a resort, not a camera, not a wait. Present in accuracy_metrics; never publishable. observations produced by a superseded wait conversion — live_crowd rows whose model_regime shows they came from the original naive calc or the M/D/1 formula deleted on 2026-04-04. They measure something the system no longer computes. All fall before the window start; inside the window this rule costs exactly one frame. dates outside 2026-04-05 to 2026-06-10 — Build period before, and off-season replays after. Itemised in scoring_window.excluded_by_the_window.
- Generated
- read-only recompute from predictions.hourly_predictions / predictions.live_corrected_hourly vs live_crowd + crowd_reports; ground-truth rule and pooling reproduced from engine/Backend/accuracy_metrics.py Generated 2026-08-27.
- MethodologyHow these figures are produced, and what we do not filter.
- Trust CenterEverything else we publish about ourselves.