Methodology
How We Research
Every growing number on GardenTrack traces to a named source, usually a university extension publication, and carries a tier saying how well evidenced it is. Before a batch merges, a checker re-fetches those sources and compares them against what we wrote. We publish what it found and the gaps we have not filled, because a number you cannot check is one you should not trust.
Most gardening sites ask you to trust them. This page is the other thing: the working out. How a number gets onto a plant page, what has to be true before it does, who checks it afterwards, what we found when they did — and the list of things we still do not know. If a figure appears anywhere on this site, its provenance starts here.
The four rules
There is no credentialed horticulturalist on this site, and inventing one would be the easiest lie in the category. What we do instead is make every number checkable. That comes down to four rules, and everything below is how they are enforced.
- A value names its source. Not a bibliography at the bottom of the page — the individual field carries the publication it came from. Days to germination for tomato cites a specific extension page, and that page is one click away.
- A value carries its confidence. “We read this in a university bulletin” and “this is a sensible default” are different claims, and they are labeled differently rather than blended into one confident voice.
- Somebody else checks it. Whoever sources a batch does not sign it off. An independent pass re-opens the original sources and grades the work.
- The gaps get published. A missing answer is reported as missing. It never gets filled with a plausible guess, and it never gets quietly dropped so the page looks complete.
What the audits found
This is the part most sites would not print. Every batch of plant data gets a checking pass before it merges — the sources fetched again and compared against what we wrote, by an automated checker that did not do the original sourcing. Across 260 batches, that pass has re-checked 12,315 values and recorded 191 problems — a finding rate of about 1.55%. Each one carries a written resolution in the batch record.
Being exact about that, because it is the sort of claim worth being exact about: 260 batches carry an audit record, and 126 of those name the independent checker that ran it. The rest record the count and the findings without naming who checked. These are automated checks re-fetching the cited pages, not a human proofreader — we would rather say so than let "independent review" do work it has not earned.
| What went wrong | Times | What it means |
|---|---|---|
| The page we cited did not say it | 160 | The value may well have been right, but the source we pointed at did not support it. Every one was either re-sourced or dropped to an uncited default. |
| The value did not match the source | 26 | The source said one thing and our field said another. |
| We said “does not apply” when it does | 3 | A field marked irrelevant for a crop that genuinely has one — summer squash marked as never direct-sown, for instance, when it usually is. |
| The arithmetic or the unit slipped | 2 | A midpoint rounded the wrong way, or a number recorded in the wrong unit. |
The shape of that table is the most useful thing we have learned about our own work. 160 of the 191 are miscites: the value itself was often defensible, and the page we pointed at simply did not say it. That is what careful sourcing fails at when you do a lot of it, and it is invisible from outside — nobody can catch it without re-opening every source.
The other 31 are the ones that matter, and some were genuinely wrong. A scuppernong grape had been written down as self-fertile when Clemson lists it as female-flowered and needing a pollinizer — plant that alone and it flowers and sets nothing, for years. A raspberry had a 5–10 gallon container against the cited Oregon State bulletin's 20–30. One plant height was marked source-verified on the strength of a quote that is not in the source at all.
The raspberry and the currant are both recorded as fixed before their batch merged. The grape's record does not say either way, so we are not going to tell you it never reached a page. That distinction is the sort of thing that gets rounded off in a sentence like this one, and rounding it off is how a page like this stops being worth reading.
We would rather show you the list than round it into the tidier number, because the tidier number is the one you should trust least. "All our mistakes were bookkeeping" is a claim to discount on sight. "We wrote down a grape that cannot fruit on its own, and the check caught it" is one you can actually weigh.
A finding is one recorded audit note, not one wrong value — a single note can cover several. One batch-211 finding covers nine candidate values rejected in one pass. Counts are generated from the batch ledgers, not typed by hand.
Every field says how well it is evidenced
A plant record is not one thing we either know or do not. It is dozens of separate claims at different strengths, and flattening them into a single confident voice is how gardening content usually goes wrong. So each field carries its own tier:
| Tier | Fields | What it means |
|---|---|---|
| Verified | 1,951 | We opened the source and it states this value. |
| Reviewed | 1,241 | Checked by a human against standard horticultural practice. |
| Cited | 1,024 | A source is attached and the value is consistent with it. |
| Estimated | 686 | A reasonable default. We would rather label it than hide it. |
Counting every record we publish, crops and varieties together, 79.23% of values sit at that top "verified" tier — we opened the source and it states the value. Others below it still carry citations; that figure is only the strongest rung, not everything with a source behind it. Whatever a value's tier, it is labeled rather than disguised.
What we do not have
Every plant site implies complete coverage. Ours is not complete, and here is the shape of the hole. Of 39 core fields, only 17 are filled for all 152 crops. These are the thinnest:
| Field | Crops we hold it for | Share |
|---|---|---|
| container min depth | 7 | 5% |
| human toxicity | 43 | 28% |
| succession interval days | 48 | 32% |
| days to maturity from transplant | 63 | 41% |
| pet toxicity | 66 | 43% |
| start indoors weeks before last frost | 78 | 51% |
| hardiness zone max | 85 | 56% |
| hardiness zone min | 85 | 56% |
The one at the top of that table matters more than the rest, so read it plainly: we hold human toxicity data for 7 of 152 crops, and pet toxicity for 66. Do not use this site to decide whether a plant is safe to eat, or safe around a child or an animal. A blank is a blank, not a clean bill of health. For that question go to your poison control center or your vet, who will give you an answer we cannot.
The others are ordinary backlog. Where a field is missing, the plant page says it is missing rather than printing something plausible.
What has to be true before a number publishes
Data is built in small batches — one crop or one field across many crops — and each batch runs the same path. Nothing skips a step because it is obviously fine.
- Source it. Fetch the actual page, quote the actual sentence. Not a summary of a summary.
- Record it with its provenance. Value, unit, source, tier — together, in one record.
- Run the gates. 14 automated checks, described below. A block stops the batch.
- Independent audit. Someone re-opens the sources and grades the batch. Findings go in the ledger whether they are flattering or not.
- Ledger and verdict. The batch gets a written verdict, and the audit findings stay on the record permanently. The table above is that record, added up.
The 14 automated gates are the mechanical floor: they check the schema, the field shapes, the source registry, licenses, provenance, identifiers, cross-references, and the safety fence on anything health-related. A value whose source is missing, unlicensed or not in the registry is a block, and a block does not produce a warning — the build stops, deletes its own output and emits nothing. So there is no version of this dataset that exists having failed them; that is what "the mechanical floor" means here, rather than a score we passed.
What the gates cannot check is whether a cited page actually says what we claim. No schema check can read English. That is what the source re-fetching is for, and it is where every one of the 191 findings above came from.
What is in the dataset
| What | How many |
|---|---|
| Crops published | 152 |
| Cultivar records | 3,202 |
| Core fields per crop | 39 |
| Growth-stage rows | 852 across 152 crops |
| Sources registered | 3,728 across 730 organizations |
| Batches built and audited | 260 |
Growth-stage rows are the unusual one. For each crop we record what the plant should be doing and when — germination, true leaves, flowering, harvest — with the signs you would actually see and the day it is late. That is the layer nobody else publishes, and it is what lets the app say your beans should have been up nine days ago.
Where the sources come from
A single “3,728 sources” headline would flatter us and mislead you, because those sources are not one kind of thing. Here is the honest split:
| Kind | Sources | What it carries |
|---|---|---|
| University extension services | 1,023 | The backbone. Land-grant university publications — the same ones a county extension agent would hand you. |
| USDA and other government records | 14 | Hardiness zones, plant profiles, official variety records. |
| Peer-reviewed research | 6 | Used where extension guidance is thin or disputed. |
| Seed catalog listings | 2,368 | Only for variety-level facts — days to maturity, color, habit. A catalog is the primary record of what a cultivar is, and a poor source for how to grow it. |
The full ledger — every source, its license, and what we are and are not allowed to do with it — is on the data sources page.
The deep dives
Some numbers need more than a paragraph. These are the full write-ups, method and measured error and all:
- How we compute frost dates — the 26-year Daymet surface, the elevation correction, and the national validation against 7,006 held-out NOAA stations, including the accuracy bar we set, missed, and had to change.
- Data sources and licenses — every dataset behind the site and the app, what it powers, and the terms it is used under.
- The frost-station extract — 7,305 stations of NOAA freeze normals, cleaned into one documented CSV and released free.
- Checks that pass on nothing — five times in two days our own automated checks passed while examining nothing at all, including one whose safeguard was defeated by the exact mistake it existed to catch. What broke, and why a green tick is only evidence if you have watched that check fail.
If we have got something wrong
We would rather hear it than not. There is no form to fill in and no account needed — email hello@gardentrack.app with the page and what is wrong, and it goes on the record the same way an audit finding does. If it changes a published number, the page changes and so does the note underneath it.
This is not politeness. On a site with no credentialed reviewer, a working correction loop is the actual quality mechanism — the audits catch what a second reader can catch, and you catch what only a gardener standing in the bed can.
The same applies to the parts we build in public. The build log carries the mistakes as they happened, including a frost warning that could never fire and three separate confidently-wrong answers about what you can plant in August.
FAQ
Where does GardenTrack get its plant data?
From named, individually recorded sources — 3,728 of them across 730 organizations. University extension publications carry the growing advice (1,023 sources); seed catalogs carry variety-level facts like days to maturity (2,368). Every field on a plant page records which source it came from, so you can check it rather than take our word.
How do you check the data is right?
Before a batch merges, an automated checker fetches the cited pages again and compares them against what we wrote — it did not do the original sourcing, so it is not marking its own homework. Across 260 batches, 12,315 values have been re-checked and 191 problems recorded, each with a written resolution. Most were citations that did not support the value; a smaller number were genuinely wrong, and this page names some of those rather than hiding them.
What is a confidence tier?
A label on each individual field saying how well evidenced it is. “Verified” means we read the source and it says exactly this. “Cited” means a source is attached. “Reviewed” means a human checked it against general horticultural knowledge. “Estimated” means it is a reasonable default and we are telling you so rather than dressing it up.
Is GardenTrack written by a qualified horticulturalist?
No, and we will not pretend otherwise. There is no Master Gardener signing off on these pages, and the checking described here is automated rather than a person re-reading every line. What we offer instead is that every value is traceable to a published source you can open yourself, that the checks are run and their findings published including the embarrassing ones, and that corrections are public. For anything where being wrong costs you a season, check your local extension office too — they know your county and we do not.
What data do you not have?
Plenty, and we list it below rather than hiding it. Of 39 core fields, 17 are filled for all 152 crops and the rest are not. Toxicity is the thinnest: we hold human toxicity for 7 crops. Do not rely on us for whether a plant is safe to eat or safe around a pet.
All of this exists so the free tools can be trusted: the frost date lookup, the companion planting chart with its myths labeled, and the spacing calculator. The plant database is the same data, crop by crop.
The numbers on this page are generated from the dataset itself every time the site builds, so they cannot drift from what is actually in it. Dataset version 2026.07.0.