Skip to content
Andrews Dean
← Field Notes
June 17, 20263 minShipping AI you can trust

The demo that lied to me

An AI improvement isn't real until it beats the noise floor: the amount your system wobbles on its own, with nothing changed.

The noise floor, visualised
0%5%10%Same input, run 5×wobbled this far, alone“Improved” prompt, +2%inside the noise floor — could be luck

My AI read the receipt and nailed every field. Vendor, date, total. Perfect.

Old me calls that a win and moves on. Instead I ran the same receipt four more times, with nothing changed. Same file, same prompt, same model.

One of the four came back different. Not wrong. It just used the shop's full legal name instead of the short one on top. But different, with nothing changed at all.

The noise floor

That gap has a name I can't un-see now. The noise floor.

AI answers wobble on their own, even when your inputs are identical. Picture a bathroom scale that shows a slightly different weight each time you step on it. Now imagine trying to prove your new diet works using only that scale. That's the trap.

Here's the part worth writing down: if you don't know how much your system wobbles by itself, you can never tell "my change helped" apart from "I got lucky this run."

A new prompt that lifts accuracy by 2% means nothing if the natural wobble is 5%. You didn't improve anything. You rolled the dice and liked the number.

Every change has to clear the line

So now, before I trust any AI improvement, it has to clear the noise floor first. I measure how much the system varies on its own, and any change has to beat that margin to count. If it can't beat the randomness, it isn't real yet.

In practice that means running the same batch several times before I run it once with a change. The wobble has to be measured on its own, with nothing touched, or there's no line to clear against. It's an extra step nobody asks for, and it's the step every roadmap tries to skip.

This is the annoying grown-up version of a rule you already know from school science. One good result proves nothing. "It worked when I tried it" is a feeling, not a finding.

The demo will always flatter you a little. Your whole job is to measure by how much.

So let me ask you. What's your team's noise floor? Because if you don't know it, every "the new version is better" is a coin flip in a lab coat.

aievaluationproduct-managementmetrics

Contact

Let's talk.

If you're building an AI-first product org — or you need someone who can take a vague mandate and return a shipped, adopted product — I'd like to hear about it.

hello@andrewsdean.com · Noida (Delhi NCR), India