4 min read ·
Your 99.9% is fifty people
Fifty students are not a rounding error. Move 50,000 student records at 99.9% accuracy and about fifty of them come out wrong. A percentage with no method behind it is decoration. Data is checked field by field against its source, or it isn't checked at all.
I know the arithmetic because the number is mine. From January 2023 to May 2024 I was a graduate assistant at the University of Illinois Chicago, and one of my jobs was moving more than 50,000 student records with Python scripts. They came out at 99.9%, verified field by field against the source.
Anyone can produce the percentage. Pull 1,000 rows at random, read them by eye and count the bad ones. With fifty bad records hidden among 50,000, you will find one on average, and you can write 99.9% as well, after looking at 2% of the table. More than a third of the time you find none at all.
A full compare of all 50,000 rows prints the same 99.9%, plus fifty names. A number that cannot tell you whether it came from the spot check or the compare is decorative accuracy. It has no denominator you can check and no list of the rows that failed, and it never says how anyone decided that a row had failed. It passes for rigor because it carries a decimal point, and you will find it on migration sign-offs, vendor data sheets, status reports and resumes. Without the seven words that follow it, my own 99.9% is decorative accuracy too. Those words promise a list, which means you get to ask me for mine.
At zero decimal places, 99.9% prints as 100%. Fifty people, rounded to nobody.
The defense arrives as a proverb: "Perfect is the enemy of good."
Nobody asked for perfect. I'm asking for the list. A field-by-field compare ends in rows, and each row is a record, a field, the value in the source and the value in the target. Turn those rows into a percentage and you throw away every name you had.
Uptime is the one place where I'll take a percentage, and 99.9% is a fair target there. That allows 8.76 hours of downtime a year, and the hours are interchangeable enough that a team can budget them and spend them. Records don't average out like that. The fiftieth student's record isn't 99.9% right. It's wrong, and so is everything downstream that reads it. Try telling the man at an empty baggage carousel that 99.9% of bags arrived.
Validation is boring. So was the cause of the biggest outage of the year. On July 19 a CrowdStrike content update crashed about 8.5 million Windows machines, by Microsoft's estimate. CrowdStrike's own root cause analysis traces it to two numbers. A template type defined 21 input fields. The sensor code supplied 20. The interpreter that read them had no runtime bounds check, and the Content Validator, the piece that exists to stop bad content before it ships, had a logic error of its own. Two of the fixes that followed were a bounds check and a check on the number of inputs.
A compare on any migration is the same kind of code, a loop over a key with an equality test inside it, the sort of thing a first-year student writes in an afternoon. The hard part is deciding what "equal" means for each field, whether a trailing space counts as a difference and whether 01/05/2023 is in January or in May, and having that argument before the cutover instead of after it.
Decorative accuracy survives because it is comfortable. The list of failed rows costs one afternoon of code, and a percentage is what you report when you would rather not be asked for it.
When the next sign-off says 99.9%, skip the congratulations.
Which fifty?
