# Forecast Accuracy Scoreboard · LiftLine AI

> The LiftLine AI forecast scoreboard: error by predicted crowd level, per lift, in minutes and in people, with the coverage denominator and the known weaknesses stated.

- Canonical human page: https://liftlineai.app/trust/accuracy/
- AI-readable edition: https://liftlineai.app/trust/accuracy/index.html.md
- Publisher: LiftLine AI
- Last verified: August 27, 2026

## What this page publishes

**Paused for the off-season.** Scoring resumes as resorts open.

This is the scoreboard. It covers one mountain, one quiet spring, and a small fraction of the forecasts we issued. Read the first table before the headline number — the headline is carried almost entirely by the hours when nothing was happening.

## When we said it would be busy, we were wrong

Forecasts are grouped by what the model itself said the crowd level would be, then scored against what the cameras and skier reports saw. This is the table we would want to read about someone else's forecast, so it goes first.

**Day-of forecast, grouped by the crowd level the model predicted**
| We said | Hours | Average error | Median error | p90 error | Bias | The line actually ran | We said it would run |
|---|---|---|---|---|---|---|---|
| Quiet | 425 | 1.85 min | 1 min | 5 min | +1.56 min | 1.94 min | 3.51 min |
| Moderate | 95 | 9.85 min | 10 min | 15 min | +9.85 min | 2.73 min | 12.58 min |
| Busy | 34 | 17.47 min | 17.5 min | 25.7 min | +17.47 min | 2.41 min | 19.88 min |

When we say a line is quiet we are usually right. When we say it is busy we are usually wrong, and wrong by about eighteen minutes — on exactly the hours a rider would change their plan over.

**The worst case, stated plainly.** On the 34 hours we called busy, we predicted 19.88 min and the line ran 2.41 min. That is a 17.47 min average error on exactly the hours a rider would act on, and it is our worst case.

## The pooled number, in context

Scored against camera-derived waits at a California ski area, the day-of forecast has a mean absolute error of 4.18 minutes across 554 forecast-hours over 100 lift-days at a California ski area, April 5 to June 10, 2026. Over those same hours the average line was 2.11 minutes long, so most of that error is us over-calling a quiet lift. It covers 12.7 percent of the forecasts issued in that window; the rest were never observed and so were never scored.

- **4.18** — minutes of average forecast error
- **2.11** — minutes — the average line those forecasts were measured against
- **+3.96** — minutes of bias — 95% of the error is one-directional over-prediction
- **12.7%** — of issued forecast-hours could be scored (554 of 4,346)

### The window, and the number without it

A lift-hour is scored only when (a) every camera observation behind it was produced by the wait conversion currently in production, and (b) it falls inside the period of continuous scheduled collection.

The published figure has to answer one question: how well does the system that exists today forecast? Observations produced by a conversion that has since been deleted cannot answer it, and neither can days on which the collector was not running.

**The published figure, and the same forecast with no window applied**
|  | Lift-hours | Lift-days | Average error | Bias | The line actually ran |
|---|---|---|---|---|---|
| Scored window, 2026-04-05 to 2026-06-10 | 554 | 100 | 4.18 min | +3.96 min | 2.11 min |
| Every scored hour on record, 2026-03-19 to 2026-08-23 | 679 | 129 | 5.34 min | +5 min | 2.44 min |

The whole scored record with no window applied, published so the windowed figure can be checked against what it was cut from. This is NOT the headline. It mixes three ground-truth conversions and counts 2026-05-08 three times.

**The whole record, month by month**
| Month | Lift-hours | Average error | Bias | In the window |
|---|---|---|---|---|
| 2026-03 | 29 | 20.69 min | +20.62 min | no |
| 2026-04 | 546 | 4.15 min | +3.74 min | partly — 2026-04-01 to 04-04 excluded |
| 2026-05 | 44 | 7.25 min | +7.21 min | yes |
| 2026-06 | 3 | 1 min | +1 min | yes |
| 2026-08 | 57 | 7.68 min | +7.61 min | no |

The window removes 125 scored hours, taking the record from 679 hours to 554. This is every hour it removes, and the rule that removes it.

- **2026-03-19 to 2026-03-31:** 29 scored hours. Ground truth produced by the pre-Q/mu conversion; only two of the five scored cameras existed; no scheduler. This is the build period. This block is the single largest mover of the headline. It is removed on the conversion-change rule, which was fixed before the per-month figures were looked at.
- **2026-04-01 to 2026-04-04:** 39 scored hours. Ground truth still produced by the M/D/1 formula (deleted 2026-04-04 12:11) or by the pre-model path, and the camera set was still changing. Removing these four days makes the headline WORSE than a plain 'drop March' cut would (4.18 against 4.25). They are removed because the rule says so.
- **after 2026-06-10:** 57 scored hours. Outside continuous collection — the scheduler had been off for two months. These are off-season dev runs.

**What the headline would have been at other start dates**
| Start date | What supports it | Lift-hours | Average error | Bias |
|---|---|---|---|---|
| 2026-04-01 | calendar April, no event behind it | 565 | 4.25 min | +4 min |
| 2026-04-05 | CHOSEN — first whole day on the current conversion with all scored cameras live | 554 | 4.18 min | +3.96 min |
| 2026-04-06 | rejected — no event supports it; it merely drops a bad day (2026-04-05 scored 14.61) | 521 | 3.52 min | +3.29 min |

Published so the boundary can be second-guessed. The chosen start is 2026-04-05 because that is where the evidence lands. It is NOT the start date that produces the best number — starting one day later would improve the headline by 0.66 minutes, and that day was kept.

**We found this while drawing the window, and it corrected our own numbers.** Every off-season scored date is a byte-identical replay of an in-season date. live_crowd frames for Broadway (47 frames) and Discovery (45) are byte-identical across 2026-05-08, 2026-08-18 and 2026-08-23: same counts, same waits, same local times-of-day, same sums. In the all-data figure, 2026-05-08 is counted three times. Excluding only the 2026-08-23 copy moves all-data MAE 5.34 -> 5.31. What it does to the published figure: None. The window ends 2026-06-10, so it contains 2026-05-08 exactly once. The detail, with the digests, is on the methodology page.

Quoting 4.18 minutes on its own would be technically true and substantively misleading. Against a line that averaged 2.11 minutes, an error of 4.18 minutes is not precision — it is us calling a crowd that was not there.

## Both series, including the one that is not a forecast

| Series | What it is | Horizon | Lift-hours | Lift-days | Average error | Median | Bias | p90 |
|---|---|---|---|---|---|---|---|---|
| Day-of AI forecast | The AI forecast, written each morning and never touched again | Same day, issued around 01:11 local | 554 | 100 | 4.18 min | 2 min | +3.96 min | 12 min |
| Live-corrected estimate — not a forecast | The current-hour number after it has already read the cameras | None | 84 | 17 | 6.19 min | 5 min | +6.12 min | 14.7 min |
| Mountain sim, day-ahead | Not scored yet | Day ahead | 0 | 0 | — | — | — | — |
| Sim after live update | Not built yet | None | 0 | 0 | — | — | — | — |

**The second row is not a forecast.** The live-corrected estimate is graded against the same camera counts it just finished reading. That is not forecasting skill and we do not present it as any. It is worse than the forecast it corrects — 6.19 min against 4.18 min — and it is here because leaving it out would be the more misleading choice. On the 84 hours where both exist, correcting changed 12 of them and improved the average by 0.42 minutes.

## How little of this gets scored

A forecast-hour can be scored only where a camera or a skier saw that lift in that hour. In the published window we issued 4,346 lift-hours of forecast and could check 554 of them — 12.7 percent. Across the whole record 6 lift lines at one resort carry a camera; 5 of them recorded frames inside this window, and 9 lifts were scored — because a lift can be scored from skier reports with no camera on it at all. Every other lift in the app is a forecast nobody has ever checked.

|  | Issued | Scoreable | Coverage |
|---|---|---|---|
| Forecast-hours | 4,346 | 554 | 12.7% |
| Lift-days | 486 | 100 | 20.6% |
| Lifts | 16 | 9 | — |

More lifts were scored (9) than carry a camera (5) because four more were scored on a handful of hours from skier reports alone.

**12.6 percent of scored hours are hindcasts, not forecasts.** 70 of 554 scored forecast-hours come from prediction rows that were written later in the day than the hour they describe. Those are hindcasts sitting inside a forecast number. We disclose them rather than remove them, because removing them would be a filter we chose after seeing the result.

## Per lift

Only five lifts have enough scored hours to say anything about.

| Lift | Hours | Days | Average error | Median | p90 | Bias | Line ran | We said |
|---|---|---|---|---|---|---|---|---|
| a measured lift | 149 | 26 | 5.62 min | 3 min | 14 min | +5.51 min | 1.92 min | 7.43 min |
| Discovery Chair | 147 | 23 | 3.39 min | 3 min | 7 min | +2.84 min | 4.56 min | 7.4 min |
| Canyon Express | 97 | 17 | 5.44 min | 1 min | 16.8 min | +5.38 min | 0.91 min | 6.29 min |
| Village Gondola | 87 | 16 | 3.62 min | 1 min | 12 min | +3.62 min | 0 min | 3.62 min |
| Face Lift Express | 68 | 12 | 1.9 min | 1 min | 5 min | +1.66 min | 1.68 min | 3.34 min |

Four more lifts — a measured lift, a measured lift, Roller Coaster Express, a measured lift — have between 1 and 2 scored hours each. That is not a sample, so we are not printing a number next to them.

## Is the forecast wrong, or is the thing measuring it wrong?

Ground truth is not measured. Cameras count heads; a queue model divides by an asserted service rate to get minutes. If that service rate is too generous, the derived 'actual' is too low, and a correct forecast would look like constant over-prediction. The observed bias is +3.96 and remarkably one-directional, which is the signature a biased ruler would leave.

**The answer.** The ruler IS biased, and in the direction that flatters this objection — but it cannot account for most of the gap. Roughly a third of the bias is attributable to a plausible service-rate error. The remaining two thirds is forecast error.

**Every scored hour re-derived under a more generous service rate**
| Service rate too generous by | Average error | Bias | The line would have run |
|---|---|---|---|
| 1x | 3.567 min | +3.347 min | 2.073 min |
| 1.5x | 3.414 min | +2.304 min | 3.116 min |
| 2x | 3.739 min | +1.218 min | 4.202 min |
| 2.6x | 4.447 min | +0.004 min | 5.416 min |
| 3x | 4.98 min | -0.855 min | 6.275 min |

The service rate that best FITS the data is about 1.5x too generous, which is consistent with the Bayesian learner having already cut it 29-47% where it had enough drain events to learn. That correction removes about 31% of the bias. Erasing the bias entirely needs 2.6x, which is a different claim and a much larger one.

**What erasing the bias entirely would mean physically**
| Lift | Nameplate | Rate we use | Rate needed to clear us |
|---|---|---|---|
| a measured lift | 3,000/hr | 905/hr (30.2%) | 348/hr (11.6%) |
| Discovery Chair | 2,400/hr | 448/hr (18.7%) | 172/hr (7.2%) |
| Canyon Express | 3,000/hr | 1,055/hr (35.2%) | 406/hr (13.5%) |
| Village Gondola | 3,600/hr | 1,924/hr (53.4%) | 740/hr (20.6%) |
| Face Lift Express | 2,400/hr | 971/hr (40.5%) | 373/hr (15.6%) |

The rates already in use imply the lifts run at 19-53% of nameplate while a line exists, which is low but arguable. The rates needed to erase the bias imply 7-21% — a detachable six-pack loading 348 people an hour with a queue in front of it. That is not a service rate, it is a broken lift.

### Hours the ruler cannot be blamed for

Isolate in-window hours in which EVERY camera frame saw 10 or fewer people in the queue. On those hours the derived wait is 0 because of the zero-threshold rule, not because of a division — no service-rate error can change them. Compare the forecast's bias there against its bias everywhere else.

|  | Hours | Our bias |
|---|---|---|
| Hours the service rate could explain | 448 | +4.14 min |
| Hours it cannot touch | 95 | +3.52 min |

On hours where the camera never saw more than about five people in line, the forecast still said 2.81 minutes, and it still says so after every correction the ruler could absorb. Twelve of those 91 hours were called moderate or busy. The over-prediction on hours the ruler cannot explain is 85% as large as on hours it could. That is not the signature of a biased ruler.

**Granting every correction at once.** On the 91 of those hours with a fully populated service rate the camera saw an average of 3.87 people in line, and we forecast 2.81 minutes. Drop the ten-person zero rule entirely and apply the best-fitting service-rate cut as well, and the wait on those hours comes to 0.25 minutes. We are still 2.56 minutes high.

This is a bound, not a measurement. It rests on the counterfactual being computable from stored columns (it is: the conversion reproduces from live_crowd.queue_filtered and live_crowd.service_rate on 2,588 of 2,591 in-window frames, the three misses being 1-minute rounding). It cannot be closed without independently timed waits.

**What would settle it.** About 30 stopwatch-timed waits at a California ski area, spread across lifts and crowd levels. The protocol, the empty log and the fitting script already exist: engine/vision/validation/fit_conversion.py and engine/vision/validation/timed_waits.csv. Until that file has rows, the split between ruler error and forecast error is a bounded inference and should be published as one.

**The honest limit.** The zero-threshold hours are the strongest evidence here, and they are also the hours where the forecast has the least to gain from being right. A reader who wants to reject this analysis should attack the assumption that a queue of four visible people is a sub-one-minute wait. Nothing in this dataset proves that; it is a physical judgement, and the timed waits would test it.

## The same forecast, in the unit we actually measure

A camera counts heads. Minutes are inferred from heads by a queue model. So here is the same forecast scored in heads, using each lift-hour's own measured service rate.

|  | Value |
|---|---|
| Lift-hours | 515 |
| What the camera counted | 23.87 people, on average |
| What the forecast implies | 74.36 people |
| Mean absolute error | 53.9 people |
| Median absolute error | 18.9 people |
| Bias | +50.49 people |

The forecast implies a line 3.12 times longer than the one in the picture. This is the wait error rescaled rather than independent evidence — but heads are the thing we measure and minutes are the thing we guess, so this is the honest way round.

## What is wrong with our forecast

First person, no hedging, in the order that matters.

1. The forecast runs high, almost always. Bias is +3.96 minutes against a mean observed wait of 2.11 minutes. 95 percent of the average error is one-directional over-prediction, not scatter.
2. When the forecast says the line is busy, it is usually wrong. On hours the model labeled 'high', it predicted 19.9 minutes and the line ran 2.4. That is a 17.5 minute average error on exactly the hours a rider would act on.
3. The number is built on a quiet spring. Every scored hour is from 5 April to 3 June at a California ski area. The mean line was 2.1 minutes long. Nothing here says anything about a Saturday in January.
4. Almost nothing gets scored. 554 of 4,346 issued forecast-hours had any ground truth at all, which is 12.7 percent. Five lift lines at one resort have cameras. Every other lift in the app is unverified, including all 51 lifts forecast at the seven other resorts.
5. The ground truth is itself a model output. A camera counts heads; a queue model turns heads into minutes using service rates that were asserted, not measured. The best-fitting correction to those rates is about 1.5x and accounts for roughly a third of the over-prediction. It does not account for the rest.
6. Any line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29 percent of every camera frame inside the window.
7. The live corrector barely corrects. On the 84 hours where both series exist it changed 12 of them and improved the average error by 0.42 minutes, and it did that by reading the same observations it is then scored against.
8. Twelve and a half percent of the scored baseline hours come from prediction rows written after the hour they describe. Those are hindcasts sitting inside a forecast number.
9. The published figure is a windowed figure. 125 scored hours were excluded by the window rule, including the 29 worst hours on record. The unwindowed number is 5.34 and is published beside it.

## What the ground truth actually is

Our ground truth is 3,635 camera observations across 5 lift lines at a California ski area (2026-04-05 to 2026-06-07), plus 13 skier-submitted waits across 7 lifts. It is not a stopwatch.

- The median frame saw 15 people in line; the largest saw 160.
- A line of ten or fewer people is recorded as a zero-minute wait by rule. That is 29.2 percent of every frame we have.
- No stopwatch has ever been held against this model. Not once. The conversion from heads to minutes has never been checked against a timed wait.

**The conversion from heads to minutes has never been checked against a timed wait.** Waits are computed as wait_minutes = round(min(Q / mu, 60)), and 0 when Q <= 10, where the service rate starts from published nameplate capacity and is then updated from observed queue drains. Where the data could correct that starting assumption it moved by up to 47 percent. On Village Gondola it never moved at all, so that lift's "measured" wait is entirely a constant we typed in.

## Numbers we withdrew

These were published and are not true as stated. They are listed rather than deleted, because a correction that leaves no trace is not a correction.

| Withdrawn claim | Verdict | What replaced it |
|---|---|---|
| about 6 minutes of average forecast error across 138 scored lift-days | does not reproduce as stated | 5.99 minutes, the unweighted mean of the per-lift-day mae column across all 138 rows, mixing the leakage-free baseline series with the live_corrected series that consumed its own ground truth. |
| 5.12 minutes of mean absolute error, 622 forecast-hours, 19 March to 3 June | superseded by this document, not wrong | That figure was correctly computed for the window it named. It has been replaced because the window itself was calendar-shaped rather than event-shaped: it began at the first camera frame, which predates the wait conversion now in production by sixteen days. The comparable all-data figure recomputed here is 5.34 over 679 hours; the difference is 57 post-season hours the earlier pass did not include. |
| every forecast gets scored | false | 12.7 percent of issued forecast-hours were scoreable inside the published window. |
| 6,294 webcam crowd counts across 17 lift lines | does not reproduce | live_crowd holds 4,522 real camera rows across 6 a California ski area lift lines, plus 3,139 synthetic rows written by the simulator on 2026-08-17 and 2026-08-18. engine/vision/config.py defines exactly 6 LIFT_CONFIGS entries, all a California ski area. Neither 6,294 nor 17 reconciles. Flagged for the owner; this figure is outside this lane's scope to fix. |

## Scope of everything above

- **Resorts:** a California ski area — a California ski area is the only resort with any ground truth. Every other resort in the catalog has predictions and zero observations.
- **Window:** 2026-04-05 to 2026-06-10, 29 days with predictions
- **Excluded:** resort_slug = 'simulator' — Synthetic data generated by the mountain sim on 2026-08-17 and 2026-08-18 under mctd_ids 9001-9011. Not a resort, not a camera, not a wait. Present in accuracy_metrics; never publishable. observations produced by a superseded wait conversion — live_crowd rows whose model_regime shows they came from the original naive calc or the M/D/1 formula deleted on 2026-04-04. They measure something the system no longer computes. All fall before the window start; inside the window this rule costs exactly one frame. dates outside 2026-04-05 to 2026-06-10 — Build period before, and off-season replays after. Itemised in scoring_window.excluded_by_the_window.
- **Generated:** read-only recompute from predictions.hourly_predictions / predictions.live_corrected_hourly vs live_crowd + crowd_reports; ground-truth rule and pooling reproduced from engine/Backend/accuracy_metrics.py Generated 2026-08-27.

- [Methodology](/trust/methodology/): How these figures are produced, and what we do not filter.
- [Trust Center](/trust/): Everything else we publish about ourselves.

_Figures measured August 27, 2026. Data from the synthetic "simulator" resort is excluded from every number._

## Canonical identity shared across LiftLine AI

- Company: LiftLine AI — https://liftlineai.app/
- Legal person: Lucas Jaeger, an individual sole proprietor doing business as LiftLine AI. There is no separate legal entity.
- Consumer product: Snowcat by LiftLine AI — https://liftlineai.app/snowcat/
- Founder: Lucas Jaeger, founder of LiftLine AI and creator of Snowcat — https://liftlineai.app/about/
- Snowcat is made by LiftLine AI and was created by Lucas Jaeger.
- LiftLine AI is independent. It is not affiliated with, endorsed by, or sponsored by a California ski area, Alterra Mountain Company, Vail Resorts, Ikon Pass, Epic Pass, or any ski resort.
- Live predictions are paused for the off-season and return for the 2026-2027 winter.
- Public figures, sources, and exclusions: https://liftlineai.app/trust/

## Use and provenance

This public Markdown companion removes navigation and presentation markup to make the canonical page easier for language models and text-based tools to interpret. It is not served selectively by user agent and does not replace the canonical human page. Product claims should remain consistent with the canonical page.
