We’re teaching the frost map behind our free frost-date lookup to see frost pockets, and we published the experiment’s pass bar before running it. Results: five of six criteria passed, one failed. Cache Valley — the model captured 19 days of a 45-day gap; our declared line was half. Under the published rules, nothing ships. Every number below, plus attempt two.
What we promised
Short version of the setup: our free frost lookup runs on an elevation-corrected frost map — Stage 1, in our build plan's naming. It's measurably better than a ZIP lookup in terrain, and it's structurally blind to cold-air pooling: the valley floor that freezes while the bench 300 feet up stays green. The experiment to fix that — a terrain-shape model of where cold air drains and pools — is Stage 2.
Before we ran it, we pre-registered the pass bar — the gate, the pass/fail line the model had to clear. In public, dated, before the first calibration run. Verbatim, from the build plan:
"It works" = positive skill over Stage 1 and ≥ 90% pair-sign accuracy and well-calibrated bands in complex terrain, without degrading flat-terrain performance.
Translating the bar's terms: "positive skill" means measurably more accurate than Stage 1. "Pair-sign accuracy" means correctly calling which of two nearby spots — one on a draining bench, one in a pooling basin — is the colder one. "Well-calibrated bands" means the date ranges we publish keep the right share of real outcomes inside them.
Plus the two named cases: Cache Valley's 45-day bench-vs-floor gap and North Ogden vs. Eden's 62-day gap, both of which had to come out correctly ordered and meaningfully separated. Six scored criteria total.
One more thing was locked before any results existed, in the pipeline code itself: what the fuzzy words mean. "Meaningfully separated" = ordered correctly and at least 50% of the true gap. "No flat-terrain degradation" = within a quarter day. "Bands maintained" = at least 88% coverage and within 4 points of Stage 1. No re-reading the words after seeing the numbers.
And the failure clause, committed with the rest: regions that fail ship nothing. Plus the promise this post exists to keep — we publish the result either way.
The results, against that bar
What ran: a six-term terrain model — how pooled a spot is, how much sky it sees at night, slope, water, urban heat — calibrated on 66 weather stations from NOAA — the US government's weather agency — inside a box covering the Wasatch Front and Cache Valley, using how much warmer or colder each station runs than expected on calm, clear nights. R² of 0.315 — terrain shape explains about a third of the station-to-station variance, and nobody should pretend otherwise. Three of the six terms came back with weak, wrong-signed fits and fell back to standard values from the published science, per the thin-data rule fixed in the spec beforehand.
How to read the table: "held out" means the station was hidden from the model during fitting and used only for grading — a hidden-station test. MAE is mean absolute error, the average size of the miss in days. "Skill" is the fraction of Stage 1's error the new model removes. And "p10–p90 coverage" is the share of real outcomes that landed inside the tool's 10–90% date range.
One scope note, so criterion 1 reads honestly. The pre-registration named out-of-region transfer pairs — Colorado cold sinks, central Pennsylvania frost hollows, Oregon valleys. This run's held-out pairs are all inside the Wasatch calibration box, evaluated leave-pair-out: the model is graded on station pairs it never saw during fitting, but the transfer test to other regions hasn't happened and nothing here claims it. That was always the rollout's validation job, and the rollout is now on hold anyway. (Peter Sinks, the fourth named stress case, did get run — it comes back below, and not flatteringly.)
| # | Pre-registered criterion | Measured | Verdict |
|---|---|---|---|
| 1 | ≥ 90% bench/basin sign accuracy on held-out pairs | 96% across 53 enumerated pairs, leave-pair-out (100% on the 14-pair unique-station subset; Stage 1 on the same pairs: 94%) | PASS |
| 2 | Positive skill over Stage 1 (median-date MAE, held-out NOAA stations, complex terrain) | Stage 2 5.82 days vs. Stage 1 6.27 → skill 0.072 | PASS |
| 3 | Cache Valley spread of the right order (truth: 45 days, bench warmer despite sitting higher) | Stage 2 19.0 days — 42% of the truth captured; the declared line was ≥ 50%. (Stage 1: 0.0 days; ordered correctly) | FAIL |
| 4 | North Ogden ↔ Eden ordered and separated (truth: ~62 days) | Stage 2 38.5 days vs. Stage 1's 19.0 (ordered correctly, truth stations held out) | PASS |
| 5 | No flat-terrain degradation | Flat MAE Stage 2 4.79 days vs. Stage 1's 7.41 — improved | PASS |
| 6 | Band calibration maintained (~90%) | Stage 2 p10–p90 coverage 97% (Stage 1 in the same test setup: 92%) | PASS |
Gate verdict: FAIL. At least one criterion missed. Under the rule we published, that's the whole determination — there is no partial credit column.
What worked
Because a fail post that hides the wins is as dishonest as a win post that hides the misses:
- The sign test. 53 held-out bench/basin pairs — one station on a draining slope, one in a pooling basin, similar elevation, the model graded only on pairs it never trained on. It called the colder side correctly in 51 of 53: 96%, against the 90% bar. Honest context, straight from the report: Stage 1 alone already got 94% of these pairs right, because the underlying 1 km base data has quietly absorbed decades of these stations' own records. So this criterion was nearly met before Stage 2 added anything — which is why the next three numbers matter more.
- North Ogden vs. Eden. The marquee same-elevation case: two towns, a bench and a basin, a real 62-day frost-season gap between them. Stage 1 sees 19 days of it — inherited from the base data, not modeled. Stage 2, with the truth stations held out, sees 38.5 days, correctly ordered. Roughly a doubling of the captured gap, from terrain shape alone.
- Complex terrain got better. Held-out station error 6.27 → 5.82 days. A 7% cut — small, real, and in the direction the hypothesis predicted.
- Flat terrain got better too. 7.41 → 4.79 days. The criterion only demanded it not get worse. We won't pretend we predicted an improvement that size on flat ground; the measured number is the measured number.
- The uncertainty bands held. 97% of held-out station values landed inside the 10–90% band, vs. 92% for Stage 1 in the same test setup.
Plain-English version of all that: on almost every test we set, the terrain model did what the physics said it should. Cold air drains downhill at night and pools in closed low spots, and a model that reads the shape of the land can mostly tell you which side of a valley gets bitten first. That part held up.
The miss
Cache Valley is the case we named, in advance, as the headline validation. Two NOAA stations in one valley: the USU bench and the Trenton valley floor, a measured 45-day difference in frost-free season — with the bench sitting about 335 feet higher and coming out warmer. The exact pattern elevation math gets backwards, the exact pattern Stage 2 exists to fix.
Stage 2 got the order right. The bench reads warmer, the floor reads colder — Stage 1 captured zero of this gap, structurally. But the size: 19.0 days of the 45. That's 42%. The line we declared before running was 50%. Miss.
The mechanism is measured, not hand-waved. To reproduce 45 days, the model needed roughly a 4 °C bench-vs-floor temperature contrast on frost-setting nights. It produced 1.71 °C. Three reasons, in order of the failure analysis:
- The tail under-correction. The model is calibrated on calm, clear nights — when cold pooling actually happens — then scaled down by a single factor (0.655) to represent all nights, windy ones included. Physically motivated, but frost dates aren't set on average nights. They're set disproportionately on exactly the calm, clear ones. Scaling the whole distribution by the all-night ratio dilutes the correction precisely where the dates get decided.
- The base data already knows these stations. NASA's underlying 1 km data has partly absorbed USU's and Trenton's own long records, so the leftover signal the model learns from — what remains after the base data's own knowledge is subtracted — understates what an unmonitored spot would need. The model's fitted settings come out conservative by construction. Honest flip side: at truly station-less frost pockets, the real correction should be at least as large as what we measured here.
- A missing variable. The model knows how deep below the cold-air dam a spot sits. It doesn't know how high above the pool floor it sits — the within-pool height that the cold-pool literature (Yoshino's variable) says matters. The USU bench sits 58 m below the dam rim but about 100 m above the valley floor, and the model still reads it as substantially pooled. Translation: the model can tell how deep the bowl is. It can't yet tell that you're perched on a ledge inside the bowl instead of sitting at the bottom. Cold settles at the bottom.
And then there's Peter Sinks — not a gate criterion (it's a research site with no published station normals), but the stress test we owed you. It's an extreme closed sink, the coldest place in Utah. The model calls it +1.26 °C — the wrong sign entirely. It reads the coldest hole in the state as a mild, draining slope. Why: at our 100 m terrain resolution the sink's small floor barely exists — the data sees a steep slope — and the point sits outside everything the model was calibrated on. The lesson is the one worth paying for: outside its calibration range, this model doesn't degrade gracefully. It can be confidently backwards. That's why every extrapolation flag in the design widens the uncertainty band instead of trusting the point value — in this pilot box, 63.1% of grid cells would have served a widened band, never a moved number.
The call: nothing ships
Per the failure clause we committed before the data existed: regions that fail ship nothing. The tool keeps serving the current elevation-corrected surface. No beta toggle, no "experimental" layer — the cold-air correction doesn't appear anywhere in the product. The planned regional rollout is on hold, and the transfer regions — Colorado, Pennsylvania, Oregon — stay untouched until the box containing my own valley clears its own gate.
The re-readings on offer here are genuinely tempting. Five of six. 42% is nearly half. Stage 1 captured zero of Cache Valley and we captured 19 days — in most shops that's a launch post. Every one of those statements is true. None of them is the test we published. The whole point of writing the bar down first is that "nearly" isn't a category the gate recognizes.
Fair question, though, because we have form here: we already moved one bar in public, when Stage 1's original target turned out to be arithmetically unreachable for any model of its class. Why not amend this one too? Because this bar isn't mis-specified — it's missed. Capturing half of a real, measured 45-day gap is exactly what a terrain model that understands this valley has to do; there's no arithmetic making it impossible, and the failure analysis names the specific missing pieces. A bar you amend is one the physics can't meet. A bar you missed is one the model didn't meet yet. Different things, handled differently.
The deeper reason the gate is strict: a terrain correction that's wrong-sized — or, like Peter Sinks, wrong-signed — tells the gardener in the basin they're safer than they are. That failure mode costs someone a crop. It's banned at the gate, not caught after.
Attempt 2, pre-registered here
The failure analysis points at two specific changes. We're committing to both now, before either has produced a number:
- Add the within-pool height term. Height above the cold pool's floor — the Yoshino variable the Cache Valley miss points at. The model gains the ability to tell a ledge partway up the bowl from the bottom of the bowl.
- Replace the single scale-down factor with a tail-aware correction. Instead of shrinking the terrain effect by one all-nights number (the 0.655 above), apply it where frost dates actually get set — weighting the correction toward the calm-clear-night end of the distribution (for the technically inclined: shifting radiative-frequency-weighted quantiles rather than the whole distribution).
The rules for attempt 2 are the same rules: the same six criteria, the same pre-fixed operationalizations (the exact numeric meanings of the fuzzy words, locked before attempt 1), the same hidden-station design, the same two named cases. No new metric gets invented after the fact, and no criterion gets quietly retired. If attempt 2 clears the bar, the correction starts shipping region by region, only where it passes, with the confidence indicator marking the difference. If it fails, you read those numbers here, at this level of detail, and the tool still doesn't change.
Further out, this is still the data layer the app's frost alert is planned to sit on — still to come, watching the live forecast against your garden's own odds. That remains roadmap, not a feature; nothing about it exists today, and nothing in this post moves it closer except the part where the underlying model gets more honest.
Why we do it this way
This post was nearly titled "Our terrain model doubled the frost-pocket signal." Every number under that headline would have been true. That's the thing about grading your own experiment after it runs — the flattering version is always available, and it's always defensible.
The bar published first is what makes this paragraph mean anything. When we eventually tell you a correction works, it'll be because it cleared a line we couldn't move, drawn before we knew whether we'd like the answer. Today the answer is: not yet, 5 of 6, and here is exactly why.
The numbers from attempt 2 land here when they're real — either way.
The frost lookup is free, no signup — still serving the elevation-corrected surface, still honest about what it can't see. The full method, this experiment included, lives on the methodology page.
The app all this is being built for remembers your garden and tells you what to do and when — and the beta is opening now. Join the beta if you want in.