Every AI vendor case study leads with a percentage. Sixty eight percent fewer launch hours. Forty percent faster delivery. Three times the output.
The number is usually true. It is also almost always useless to you, and the reason is arithmetic rather than dishonesty.
A percentage measures a bottleneck, not a tool
When a company reports that a model cut a process by 68%, they have not measured the capability of the model. They have measured how blocked they were.
Take the common version. A team has a fixed launch deadline and its design resources are committed elsewhere. A model absorbs the production work that was waiting on those designers, and weeks compress into days.
That 68% is a measurement of the queue in front of the design team. Remove the same constraint with a contractor, a template library, or a reshuffled roadmap, and you get a similar number.
Any aerospace engineer recognises this instantly. It is Amdahl's law wearing a marketing headline. Your total improvement is capped by the fraction of the process you actually touched.
If production work is 70% of your launch effort, automating it well produces a dramatic figure. If production is 15% of your effort because your real constraint is legal review or a slow approval chain, the same tool produces a rounding error.
The tool did not change between those two companies. The shape of the work did.
The denominator is never published
Sixty eight percent of what, exactly.
Launch hours. Which hours, measured by whom, over what baseline period, using what definition of a launch.
Case studies publish the numerator and bury the denominator, because the denominator is where the embarrassment lives. A baseline captured during a crunch quarter makes any subsequent improvement look structural rather than seasonal.
In every engagement I have run, roughly half the claimed AI savings evaporate the moment you ask what the process cost before anyone was watching it. Not because people lied, but because nobody had measured the old process at all.
You cannot report a saving against a number you never took. So teams reconstruct the baseline from memory, and memory is generous about how long things used to take.
There is a second distortion in every published figure. The team in a case study is, by definition, the team that succeeded, and the study is written after the outcome is known.
You are reading a survivor. For every published 68%, there is an unpublished set of companies that ran a similar project, measured 4%, and never spoke to the vendor's marketing team again.
This is also why so many pilots produce a spectacular result and then fail to change anything at the company level. The saving was real inside the pilot and imaginary inside the P&L. I made the longer version of this argument in why most AI pilots die in the middle.
Three numbers worth more than any case study
If you want a figure you can defend to a board, measure these instead.
Cycle time on one named workflow. Pick a single process with a clear start and end. Measure it for four weeks before you introduce anything. Four weeks, not four days, because the variance in most creative and technical work is weekly rather than daily.
Rework rate. The percentage of output that comes back for a second pass. Nobody publishes this, and it is where machine assisted work quietly loses the gains it made on the first pass. Faster first drafts with a higher rework rate can be slower end to end.
Cost per shipped unit, not cost per generated unit. Generation is cheap and getting cheaper. Shipping requires review, correction, approval, and someone accepting responsibility. Those costs did not fall, and in several teams I have watched them rise, because reviewing machine output is a different and more tiring job than reviewing a colleague's.
Measure all three on the same workflow rather than across different ones. Teams routinely compare cycle time on the new process against rework on the old one, which flatters both and proves nothing.
Three numbers, one workflow, four weeks of baseline. That is the whole method, and most companies skip all of it because the baseline period feels like a delay rather than the measurement it actually is.
What a real pilot looks like
If you cannot state your baseline in one sentence with a number in it, you are not running a pilot. You are running a demo with a budget attached.
A real pilot names the workflow, names the owner, states the pre-existing cycle time, sets the measurement window, and defines in advance what result would cause you to stop. That last item is the one that gets removed from every pilot charter I read.
Without a stopping condition, a pilot cannot fail. It can only be extended, rebranded, or absorbed into a larger initiative where nobody has to publish a number.
The discipline here is not about scepticism toward AI. I build with these systems every week and the gains are real where the constraint is real.
The discipline is about knowing which constraint you actually have, because that determines your ceiling before you evaluate a single vendor.
The vendors are not lying to you. They are answering a precise question about their own business, and you mistook it for a question about yours.