Our free frost-date lookup can’t see frost pockets — low spots where cold air pools and frost lands weeks early. We’re about to build a terrain model to fix that. The pass bar, published first: call the colder side of station pairs it never trained on, 90% minimum; beat today’s map; don’t degrade flat ground. We publish the result either way.
The setup: what our frost map can't see, measured
This week we shipped a new engine for the free frost lookup: frost odds computed for your own ~800-foot square, corrected for its elevation, from 26 years of NASA daily data — one big precomputed frost map we call the surface. The launch post has the measured accuracy numbers — about 5 days better than a nearest-station lookup in complex terrain.
It also has a limits section. This post is about the biggest limit, and the experiment we're about to run against it.
The surface corrects for elevation. It is blind to cold-air pooling — and the stations of NOAA, the US government's weather agency, show exactly how much that blindness costs. The numbers below are NOAA's official 30-year averages, the 1991–2020 "normals":
| Case (NOAA stations, 1991–2020 normals) | Real frost-free-season gap | What our surface captures |
|---|---|---|
| Cache Valley, UT — USU bench vs. Trenton valley floor | 45 days shorter on the floor | 0 days |
| North Ogden (bench) vs. Eden (basin), UT | 62 days (186 vs. 124) | ~20 days — 19 of them inherited from NASA's underlying 1 km data, not from anything we modeled |
Those aren't estimates. We measured our own surface against the stations and published the misses.
Here's the plain-English version of why elevation math can't fix this. On a clear, calm night, cold air behaves like water. It slides downhill and pools in the flat low spots. A bench sheds it; a closed basin holds it. So the valley floor freezes while a spot 300 feet up the slope stays green. The Cache Valley bench sits about 335 feet higher than the floor — and it's the warm one. Any correction built on "higher means colder" sees that backwards. North Ogden and Eden are worse: the two towns sit at essentially the same elevation, so an elevation-only model calls them near-identical. The stations say 62 days apart.
The hypothesis
Terrain shape can predict that gap. Not elevation — shape. How flat and low a valley bottom is. How much open sky a spot sees at night. Whether cold air drains through it or dead-ends in it. Whether it sits on a bench or in a basin.
This isn't our idea. There's decades of published science saying it should work — and we want to be precise about what that literature is and isn't:
- Lundquist, Pepin & Rochford (2008) built an automated algorithm that flags cold-air-pooling terrain from slope, relative elevation, and curvature alone.
- Curtis et al. (2014) measured Sierra Nevada stations flagged by that algorithm running a mean 1.6 °C colder than their neighbors, across over a century of records.
- Chung et al. (2006) measured a 6.7 °C nighttime-minimum spread across just 73 m of elevation — terrain, not altitude, doing the work. A later refinement of the same method explained 78% of the daily-minimum error from cold-air accumulation.
- Gustavsson et al. (1998) measured site-to-site spreads up to ~10 °C on clear nights with barely any wind — and near zero once wind mixes the air. This only works on calm, clear nights. Those happen to be exactly the nights that kill plants.
The bridge from degrees to dates: near the frost thresholds, daily minimums move about 0.2–0.35 °C per day through spring and fall. So a 1.6 °C cold-pool offset is roughly 5–8 days on the calendar, and a 3 °C offset is 9–15 days. That's the right order of magnitude for the gaps in the table above.
We have one piece of our own evidence that the headroom is real: in our Nevada validation region, 36% of grid cells came back inversion-flagged — places where the temperature-vs-elevation relationship goes flat or backwards, the signature of pooling cold air. That's a marker of where the problem lives, not proof we can solve it.
And the honest counterweight, from the same literature: the tuning numbers a model learns in one region transfer poorly to another. Cold air doesn't route exactly like water — it spills and dams in ways a simple flow model misses. And no bare-earth terrain model can see a road berm or a hedge damming cold air behind your house. The physics says this should work. "Should" is not a measurement. I've planted by "should" before — timed my corn and beans exactly the way everything online said to, and the beans still had nothing to climb when they were ready. So we test.
The bar, committed before the data exists
Here's the problem with grading your own experiment after it runs: there's always some metric that flatters the result. The fix is boring and old — pre-registration. State the pass bar first, in public, then run.
So, verbatim from our build plan, the gate this model must clear before it touches the tool. (In the plan's naming: the elevation-corrected surface that's live today is Stage 1; the cold-air-drainage correction is Stage 2.)
"It works" = positive skill over Stage 1 and ≥ 90% pair-sign accuracy and well-calibrated bands in complex terrain, without degrading flat-terrain performance.
Translating the bar's terms: "positive skill" means measurably more accurate than what's live today. "Pair-sign accuracy" means correctly calling which of two nearby spots — one on a draining bench, one in a pooling basin — is the colder one. "Well-calibrated bands" means the date ranges we publish stay honest: the right share of real outcomes lands inside them.
In plain terms, four checks, all required:
- ≥ 90% bench/basin sign accuracy on held-out pairs. Matched station pairs — one on a draining bench, one in a pooling basin, similar elevation — that the model never saw during calibration: Colorado cold sinks, central Pennsylvania frost hollows, Oregon valleys, Utah's Peter Sinks. The model must call which side is colder at least 9 times in 10.
- Positive skill over the current surface. On frost-pocket pairs it must measurably beat what's live on the tool today. A tie means it doesn't ship.
- Cache Valley and North Ogden/Eden correctly separated. The two cases in the table above — the 45 days we currently capture none of, and the 62-day gap we see less than a third of — must come out correctly ordered and meaningfully separated.
- No flat-terrain degradation. A model tuned for mountain valleys must not make Kansas worse. Flat-ground accuracy holds or the whole thing fails.
And the failure clause, also committed now: regions that fail ship nothing. Anywhere the model can't show a measurable improvement on stations hidden from its calibration, the tool keeps serving the current elevation-corrected surface and the confidence indicator says so. A wrong-signed terrain correction is worse than none — it tells the gardener in the basin they're safer than they are. That failure mode costs someone a crop, so it's banned at the gate, not caught after.
The promise: we publish the result either way. If the model clears the bar, you'll read the numbers here. If it fails, you'll read those numbers here, in the same detail, and the tool won't change.
Why pre-commit? Because we've already moved one bar in public
Fair question: why not just run the experiment and report honestly after?
Because we've been on the other side of this once already. Our original accuracy bar for the current surface said: beat the nearest-station baseline by at least 50% in complex terrain. We wrote that bar before we'd measured either side of the ratio. Then two validation regions came back at 38.9% and 44.7% — and the root cause wasn't a fixable bug. Our error floor is the underlying NASA data's own ~6-day error at stations, and the baseline only degrades to ~11 days even when the nearest station is 30 km away. The arithmetic caps the ratio near 45%. No region rescues it. The bar was mis-specified, not missed.
So we changed it — in public. The old bar is struck through in the plan, not deleted. The amendment is dated, the reasoning is spelled out, and both failing pilot reports are published in full on the methodology page. The re-anchored bar is absolute instead of relative: error of 6–7 days or less in complex terrain, at least a 40% cut where stations are sparse, no flat-terrain degradation, honest probability bands. Both pilots pass it, and the tool's public claim was calibrated down to match what was actually measured.
We think that was the right call, handled the right way. But moving a bar after seeing the data is only defensible if you show every step of the work — and it's a door you don't want standing open. The way to close it is to commit the next bar before the data exists. That's this post. The Stage-2 bar above is now fixed in public, dated before the first calibration run. If it turns out to be mis-specified too, amending it will cost us another public explanation — as it should.
What this becomes if it works
If the model clears the bar, the free tool gets the terrain correction, region by region, only where it passed — with the confidence indicator marking the difference.
And it's the data layer for something further out. The GardenTrack app — now in beta — is planned to include a frost alert: it knows your garden's frost odds from this same surface, watches the live forecast, and gives you a heads-up when frost's in your local forecast. If Stage 2 works, that alert wouldn't reason from the town average. It would know whether your beds sit in the pocket or above it, and time its warnings to your odds. That's roadmap, not a feature — nothing about it exists today. But it's why this experiment is worth running carefully instead of fast.
The bug story behind the current surface is the previous post, if you want to see what our version of "carefully" looks like in practice. The experiment starts now. The numbers land here when they're real.
The frost lookup is free, no signup — honest about what it can and can't see.
The app all this is being built for remembers your garden and tells you what to do and when — and the beta is opening now. Join the beta if you want in.