We opened a backlog of 3,472 missing plant-data values, ranked it largest-block-first, and started at the top. Three separate blocks — 1,401 pairs, 40% of the total — turned out never to have been work at all. There is a mechanism behind that: a field that does not exist at species level is missing from every species, so being fake is what makes a gap look big.
Ask a gardening app how deep to sow something, and it will always answer
That is the problem. Every app and every search result gives you a number for every crop, every time, and none of them tell you which numbers somebody actually looked up and which are filler. Sow depth for a common bean is genuinely published by half a dozen universities. Sow depth for something unusual often is not published anywhere — and you still get an answer, delivered with exactly the same confidence.
We decided early that a blank is better than a guess. Our plant records hold 40-odd facts for each of 152 crops — how deep, how far apart, the temperature that damages it, how long until you can eat it — and each one has to come from a named university extension service. Where we have not found one, the slot stays empty and the page says so.
Which means we can count our own ignorance. Crops multiplied by facts, minus what is filled: 3,472 blanks. That list is the honest version of a to-do list, and it is how we decide what to go and research next.
So we sorted it biggest-first and started at the top. That turned out to be the mistake.
Three times, the thing at the front was not a job
831 pairs: the field does not apply to that kind of plant. A sowing depth on a tree you buy as a sapling. A transplant date for something you never transplant. Our records already declared which kinds each field applies to — the declaration was written down and the thing counting the gaps could not read it.
152 pairs: a field we had retired. Its own record says, in as many words, do not populate this. It was still being counted as outstanding work by a census that did not look at that flag.
418 pairs: four fields that live one level down. They describe an individual variety — this particular tomato, not tomatoes — so at the crop level they are empty by design and always will be.
That is 1,401 pairs. Forty percent of the backlog, and none of it was ever work.
The mechanism, which is the part worth stealing
Three in a row is not luck, and the reason is almost arithmetic.
A phantom gap is large because it is a phantom. A field that does not exist at the crop level is missing from every single one of the 152 crops — that is the largest a block can possibly be. There is no bigger number available. Meanwhile a real gap is always smaller, because some crops already carry a value; that is what makes it real.
So size and reality are different axes, and ranking by size is anti-correlated with ranking by real. Sorting a backlog biggest-first does not surface the most work. It surfaces the things that were never work, and it puts them at the top where somebody confident and busy will pick them up.
We had built a queue that recommended its own fakes.
What it would have cost to keep going
The biggest block left after the phantoms was disease resistance, at 152 crops. It is also the single least sourceable thing in the file: resistance is a property of a named variety, not of the crop, and our own field contract says so.
Had we worked it from the top of the queue, we would have produced a handful of resistance codes copied from seed-catalog marketing, recorded at the wrong level, for a field that cannot hold them. And the sourcing note on each one would have looked perfectly respectable. That is the part that should worry anybody who keeps a dataset: the failure would not have looked like a failure.
The fix costs one lookup, and it is not a search
Before taking the next block off the top, read that field's own contract entry: what level it belongs to, which kinds of plant it applies to, whether it has been retired. That is a lookup in a file we already have. No web searches, no sources to read, nothing to verify.
It moved 1,401 pairs out of a queue that was about to be worked by hand.
Rank by whether a thing can be sourced, not by how big it looks.
The same shape, three times, and it is the one we keep finding
In all three cases the decision had already been made and written down. Which plant kinds a field applies to. That a field was retired. Which level a field belongs to. Every one of those was recorded in prose, or in a shape the counting script could not consume.
So the fix was never to make a new decision. It was to move an existing one into a slot a program can read — a retired flag with its reason, kind tokens beside the scope note, an explicit level on the field.
The knowledge was in the repository the whole time. The instrument that needed it could not get at it. We have now met that exact shape often enough to have written it down as a rule, which is its own small admission: if a fact only exists in a sentence, no check can defend it.
The data itself, free and sourced: the plant database.
How we decide what a source is good enough to say: how we research.
The app is in beta now. Join the beta if you want in.