Trust Center

How wrong we are, and where.

The LiftLine AI forecast scoreboard: error by predicted crowd level, per lift, in minutes and in people, with the coverage denominator and the known weaknesses stated.

This is the scoreboard. It covers one mountain, one quiet spring, and a small fraction of the forecasts we issued. Read the first table before the headline number — the headline is carried almost entirely by the hours when nothing was happening.

When we said it would be busy, we were wrong

Forecasts are grouped by what the model itself said the crowd level would be, then scored against what the cameras and skier reports saw. This is the table we would want to read about someone else's forecast, so it goes first.

Day-of forecast, grouped by the crowd level the model predicted
We saidHoursAverage errorMedian errorp90 errorBiasThe line actually ranWe said it would run
Quiet4251.85 min1 min5 min+1.56 min1.94 min3.51 min
Moderate959.85 min10 min15 min+9.85 min2.73 min12.58 min
Busy3417.47 min17.5 min25.7 min+17.47 min2.41 min19.88 min

When we say a line is quiet we are usually right. When we say it is busy we are usually wrong, and wrong by about eighteen minutes — on exactly the hours a rider would change their plan over.

The pooled number, in context

Scored against camera-derived waits at Mammoth, the day-of forecast has a mean absolute error of 4.18 minutes across 554 forecast-hours over 100 lift-days at Mammoth, April 5 to June 10, 2026. Over those same hours the average line was 2.11 minutes long, so most of that error is us over-calling a quiet lift. It covers 12.7 percent of the forecasts issued in that window; the rest were never observed and so were never scored.

4.18minutes of average forecast error
2.11minutes — the average line those forecasts were measured against
+3.96minutes of bias — 95% of the error is one-directional over-prediction
12.7%of issued forecast-hours could be scored (554 of 4,346)

The window, and the number without it

A lift-hour is scored only when (a) every camera observation behind it was produced by the wait conversion currently in production, and (b) it falls inside the period of continuous scheduled collection.

The published figure has to answer one question: how well does the system that exists today forecast? Observations produced by a conversion that has since been deleted cannot answer it, and neither can days on which the collector was not running.

The published figure, and the same forecast with no window applied
Lift-hoursLift-daysAverage errorBiasThe line actually ran
Scored window, 2026-04-05 to 2026-06-105541004.18 min+3.96 min2.11 min
Every scored hour on record, 2026-03-19 to 2026-08-236791295.34 min+5 min2.44 min

The whole scored record with no window applied, published so the windowed figure can be checked against what it was cut from. This is NOT the headline. It mixes three ground-truth conversions and counts 2026-05-08 three times.

The whole record, month by month
MonthLift-hoursAverage errorBiasIn the window
2026-032920.69 min+20.62 minno
2026-045464.15 min+3.74 minpartly — 2026-04-01 to 04-04 excluded
2026-05447.25 min+7.21 minyes
2026-0631 min+1 minyes
2026-08577.68 min+7.61 minno

The window removes 125 scored hours, taking the record from 679 hours to 554. This is every hour it removes, and the rule that removes it.

2026-03-19 to 2026-03-31
29 scored hours. Ground truth produced by the pre-Q/mu conversion; only two of the five scored cameras existed; no scheduler. This is the build period. This block is the single largest mover of the headline. It is removed on the conversion-change rule, which was fixed before the per-month figures were looked at.
2026-04-01 to 2026-04-04
39 scored hours. Ground truth still produced by the M/D/1 formula (deleted 2026-04-04 12:11) or by the pre-model path, and the camera set was still changing. Removing these four days makes the headline WORSE than a plain 'drop March' cut would (4.18 against 4.25). They are removed because the rule says so.
after 2026-06-10
57 scored hours. Outside continuous collection — the scheduler had been off for two months. These are off-season dev runs.
What the headline would have been at other start dates
Start dateWhat supports itLift-hoursAverage errorBias
2026-04-01calendar April, no event behind it5654.25 min+4 min
2026-04-05CHOSEN — first whole day on the current conversion with all scored cameras live5544.18 min+3.96 min
2026-04-06rejected — no event supports it; it merely drops a bad day (2026-04-05 scored 14.61)5213.52 min+3.29 min

Published so the boundary can be second-guessed. The chosen start is 2026-04-05 because that is where the evidence lands. It is NOT the start date that produces the best number — starting one day later would improve the headline by 0.66 minutes, and that day was kept.

Quoting 4.18 minutes on its own would be technically true and substantively misleading. Against a line that averaged 2.11 minutes, an error of 4.18 minutes is not precision — it is us calling a crowd that was not there.

Both series, including the one that is not a forecast

SeriesWhat it isHorizonLift-hoursLift-daysAverage errorMedianBiasp90
Day-of AI forecastThe AI forecast, written each morning and never touched againSame day, issued around 01:11 local5541004.18 min2 min+3.96 min12 min
Live-corrected estimate — not a forecastThe current-hour number after it has already read the camerasNone84176.19 min5 min+6.12 min14.7 min
Mountain sim, day-aheadNot scored yetDay ahead00
Sim after live updateNot built yetNone00

How little of this gets scored

A forecast-hour can be scored only where a camera or a skier saw that lift in that hour. In the published window we issued 4,346 lift-hours of forecast and could check 554 of them — 12.7 percent. Across the whole record 6 lift lines at one resort carry a camera; 5 of them recorded frames inside this window, and 9 lifts were scored — because a lift can be scored from skier reports with no camera on it at all. Every other lift in the app is a forecast nobody has ever checked.

IssuedScoreableCoverage
Forecast-hours4,34655412.7%
Lift-days48610020.6%
Lifts169

More lifts were scored (9) than carry a camera (5) because four more were scored on a handful of hours from skier reports alone.

Per lift

Only five lifts have enough scored hours to say anything about.

LiftHoursDaysAverage errorMedianp90BiasLine ranWe said
Broadway Express149265.62 min3 min14 min+5.51 min1.92 min7.43 min
Discovery Chair147233.39 min3 min7 min+2.84 min4.56 min7.4 min
Canyon Express97175.44 min1 min16.8 min+5.38 min0.91 min6.29 min
Village Gondola87163.62 min1 min12 min+3.62 min0 min3.62 min
Face Lift Express68121.9 min1 min5 min+1.66 min1.68 min3.34 min

Four more lifts — Gold Rush Express, Stump Alley Express, Roller Coaster Express, Chair 23 — have between 1 and 2 scored hours each. That is not a sample, so we are not printing a number next to them.

Is the forecast wrong, or is the thing measuring it wrong?

Ground truth is not measured. Cameras count heads; a queue model divides by an asserted service rate to get minutes. If that service rate is too generous, the derived 'actual' is too low, and a correct forecast would look like constant over-prediction. The observed bias is +3.96 and remarkably one-directional, which is the signature a biased ruler would leave.

Every scored hour re-derived under a more generous service rate
Service rate too generous byAverage errorBiasThe line would have run
1x3.567 min+3.347 min2.073 min
1.5x3.414 min+2.304 min3.116 min
2x3.739 min+1.218 min4.202 min
2.6x4.447 min+0.004 min5.416 min
3x4.98 min-0.855 min6.275 min

The service rate that best FITS the data is about 1.5x too generous, which is consistent with the Bayesian learner having already cut it 29-47% where it had enough drain events to learn. That correction removes about 31% of the bias. Erasing the bias entirely needs 2.6x, which is a different claim and a much larger one.

What erasing the bias entirely would mean physically
LiftNameplateRate we useRate needed to clear us
Broadway Express3,000/hr905/hr (30.2%)348/hr (11.6%)
Discovery Chair2,400/hr448/hr (18.7%)172/hr (7.2%)
Canyon Express3,000/hr1,055/hr (35.2%)406/hr (13.5%)
Village Gondola3,600/hr1,924/hr (53.4%)740/hr (20.6%)
Face Lift Express2,400/hr971/hr (40.5%)373/hr (15.6%)

The rates already in use imply the lifts run at 19-53% of nameplate while a line exists, which is low but arguable. The rates needed to erase the bias imply 7-21% — a detachable six-pack loading 348 people an hour with a queue in front of it. That is not a service rate, it is a broken lift.

Hours the ruler cannot be blamed for

Isolate in-window hours in which EVERY camera frame saw 10 or fewer people in the queue. On those hours the derived wait is 0 because of the zero-threshold rule, not because of a division — no service-rate error can change them. Compare the forecast's bias there against its bias everywhere else.

HoursOur bias
Hours the service rate could explain448+4.14 min
Hours it cannot touch95+3.52 min

On hours where the camera never saw more than about five people in line, the forecast still said 2.81 minutes, and it still says so after every correction the ruler could absorb. Twelve of those 91 hours were called moderate or busy. The over-prediction on hours the ruler cannot explain is 85% as large as on hours it could. That is not the signature of a biased ruler.

This is a bound, not a measurement. It rests on the counterfactual being computable from stored columns (it is: the conversion reproduces from live_crowd.queue_filtered and live_crowd.service_rate on 2,588 of 2,591 in-window frames, the three misses being 1-minute rounding). It cannot be closed without independently timed waits.

The same forecast, in the unit we actually measure

A camera counts heads. Minutes are inferred from heads by a queue model. So here is the same forecast scored in heads, using each lift-hour's own measured service rate.

Value
Lift-hours515
What the camera counted23.87 people, on average
What the forecast implies74.36 people
Mean absolute error53.9 people
Median absolute error18.9 people
Bias+50.49 people

The forecast implies a line 3.12 times longer than the one in the picture. This is the wait error rescaled rather than independent evidence — but heads are the thing we measure and minutes are the thing we guess, so this is the honest way round.

What is wrong with our forecast

First person, no hedging, in the order that matters.

  1. The forecast runs high, almost always. Bias is +3.96 minutes against a mean observed wait of 2.11 minutes. 95 percent of the average error is one-directional over-prediction, not scatter.
  2. When the forecast says the line is busy, it is usually wrong. On hours the model labeled 'high', it predicted 19.9 minutes and the line ran 2.4. That is a 17.5 minute average error on exactly the hours a rider would act on.
  3. The number is built on a quiet spring. Every scored hour is from 5 April to 3 June at Mammoth. The mean line was 2.1 minutes long. Nothing here says anything about a Saturday in January.
  4. Almost nothing gets scored. 554 of 4,346 issued forecast-hours had any ground truth at all, which is 12.7 percent. Five lift lines at one resort have cameras. Every other lift in the app is unverified, including all 51 lifts forecast at the seven other resorts.
  5. The ground truth is itself a model output. A camera counts heads; a queue model turns heads into minutes using service rates that were asserted, not measured. The best-fitting correction to those rates is about 1.5x and accounts for roughly a third of the over-prediction. It does not account for the rest.
  6. Any line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29 percent of every camera frame inside the window.
  7. The live corrector barely corrects. On the 84 hours where both series exist it changed 12 of them and improved the average error by 0.42 minutes, and it did that by reading the same observations it is then scored against.
  8. Twelve and a half percent of the scored baseline hours come from prediction rows written after the hour they describe. Those are hindcasts sitting inside a forecast number.
  9. The published figure is a windowed figure. 125 scored hours were excluded by the window rule, including the 29 worst hours on record. The unwindowed number is 5.34 and is published beside it.

What the ground truth actually is

Our ground truth is 3,635 camera observations across 5 lift lines at mammoth (2026-04-05 to 2026-06-07), plus 13 skier-submitted waits across 7 lifts. It is not a stopwatch.

  • The median frame saw 15 people in line; the largest saw 160.
  • A line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29.2 percent of every frame we have.
  • No stopwatch has ever been held against this model. Not once. The conversion from heads to minutes has never been checked against a timed wait.

Numbers we withdrew

These were published and are not true as stated. They are listed rather than deleted, because a correction that leaves no trace is not a correction.

Withdrawn claimVerdictWhat replaced it
about 6 minutes of average forecast error across 138 scored lift-daysdoes not reproduce as stated5.99 minutes, the unweighted mean of the per-lift-day mae column across all 138 rows, mixing the leakage-free baseline series with the live_corrected series that consumed its own ground truth.
5.12 minutes of mean absolute error, 622 forecast-hours, 19 March to 3 Junesuperseded by this document, not wrongThat figure was correctly computed for the window it named. It has been replaced because the window itself was calendar-shaped rather than event-shaped: it began at the first camera frame, which predates the wait conversion now in production by sixteen days. The comparable all-data figure recomputed here is 5.34 over 679 hours; the difference is 57 post-season hours the earlier pass did not include.
every forecast gets scoredfalse12.7 percent of issued forecast-hours were scoreable inside the published window.
6,294 webcam crowd counts across 17 lift linesdoes not reproducelive_crowd holds 4,522 real camera rows across 6 Mammoth lift lines, plus 3,139 synthetic rows written by the simulator on 2026-08-17 and 2026-08-18. engine/vision/config.py defines exactly 6 LIFT_CONFIGS entries, all Mammoth. Neither 6,294 nor 17 reconciles. Flagged for the owner; this figure is outside this lane's scope to fix.

Scope of everything above

Resorts
mammoth — Mammoth is the only resort with any ground truth. Every other resort in the catalog has predictions and zero observations.
Window
2026-04-05 to 2026-06-10, 29 days with predictions
Excluded
resort_slug = 'simulator' — Synthetic data generated by the mountain sim on 2026-08-17 and 2026-08-18 under mctd_ids 9001-9011. Not a resort, not a camera, not a wait. Present in accuracy_metrics; never publishable. observations produced by a superseded wait conversion — live_crowd rows whose model_regime shows they came from the original naive calc or the M/D/1 formula deleted on 2026-04-04. They measure something the system no longer computes. All fall before the window start; inside the window this rule costs exactly one frame. dates outside 2026-04-05 to 2026-06-10 — Build period before, and off-season replays after. Itemised in scoring_window.excluded_by_the_window.
Generated
read-only recompute from predictions.hourly_predictions / predictions.live_corrected_hourly vs live_crowd + crowd_reports; ground-truth rule and pooling reproduced from engine/Backend/accuracy_metrics.py Generated 2026-08-27.

Figures measured August 27, 2026. Data from the synthetic "simulator" resort is excluded from every number.