Your Team Feels 20% Faster With AI. Nobody Can Prove It.

A VP of Engineering told me last month that his team was "obviously" 30% faster since rolling out AI coding assistants. I asked what the number was measuring against. He paused, then said the vibe in standup had changed. That was the evidence. Not a ticket-cycle chart, not a deploy-frequency graph, not a single number that existed before the rollout and could be compared to one after it. The vibe.
He's not an outlier. He's the median case. Companies are spending real money on AI coding tools and feeling faster, and when you ask them to show their work, the work isn't there. That's not a footnote to the AI-productivity story. It's the story. The entire category is being purchased, defended in budget meetings, and written up in case studies based on how the work feels. Not on what the output actually shows. A feeling isn't a metric, no matter how many people share it in the same standup, and no amount of shared enthusiasm turns a vibe into a number you could put in front of a CFO.
The $200-a-Month Line Item Nobody Can Justify
DX, the engineering-intelligence firm that a large share of enterprise platform teams use to run their own internal metrics, put a number on the going rate for AI coding tools this year: roughly $200 per developer per month once you stack the assistant, the agent runtime, and the usage-based overages on top of the base seat. That's not a rounding error in an engineering budget. For a 200-person engineering org, it's a $40,000-a-month line, every month, indefinitely.
Internal usage-limit data DX has published shows about 30% of developers regularly hitting their usage caps — a real signal of heavy, sustained use, not a novelty tool people tried once and abandoned. That part of the story is genuinely encouraging. People are using this constantly, by choice, without being told to.
What DX's own research is honest about, and what most vendor pitches quietly skip, is that heavy usage isn't the same evidence as verified output. You can max out a tool's usage cap every single day and still not know whether the org is shipping more, shipping better, or just shipping differently. Fewer than one in three AI-tool decision-makers, per 2026 industry survey data, can actually connect their AI spend to a measurable business outcome. Two-thirds of the people signing the checks are renewing a $200-per-seat line item on faith. Not skepticism, not caution — faith, the same category of belief you'd need to keep funding anything you can't measure and don't want to stop believing in.
Why Developers Feel Faster While Measuring Slower
Here's the finding that should be reshaping how every engineering org talks about this, and mostly isn't. A controlled study measuring experienced open-source developers working in their own repositories — not toy benchmarks, their actual codebases, the systems they know cold — found that those developers self-reported feeling about 20% faster with AI assistance turned on. Measured completion time on the same tasks came in around 19% slower.
Read that gap twice. It's not a rounding error, and it's not one bad study cherry-picked to make a point. It's a documented perception-versus-measurement split, discussed and debated inside the industry precisely because it's uncomfortable, and it should worry you more than either number alone. The developers weren't lying. They weren't badly calibrated in some obvious, dismissible way. Something about the experience of working with an AI assistant reads as speed to the person doing the work, even when the clock says otherwise. Less typing. Less blank-page friction. More forward motion in the moment. Reviewing suggestions, correcting near-misses, re-orienting after an AI takes the task somewhere you didn't intend — none of that friction registers the same way idle staring at a blank editor does. It feels like progress. The stopwatch disagrees.
This is worth sitting with because it means the single most common piece of evidence offered for AI coding ROI doesn't hold up on its own. "My team says it's faster" is exactly the kind of evidence this research suggests you shouldn't trust in isolation. Not because developers are unreliable narrators of their own lives. Because speed as experienced and speed as measured have just been shown to diverge, on real tasks, under real working conditions, for developers who know their own codebases better than anyone auditing them ever will. That's not a reason to distrust developers. It's a reason to stop treating their self-report as the whole audit, and start treating it as one input among several.
Ninety-Five Percent of Pilots Die Quietly
Vendor decks in this category tend to cite productivity gains somewhere between 30% and 55%. Independent tracking of enterprise AI pilots tells a very different story about durability: industry estimates put the failure rate for pilots to show sustained value past six months as high as roughly 95%. Not "underwhelming." Failed to demonstrate the thing they were greenlit to demonstrate, quietly, past the point where anyone in the room still remembered the original pitch deck.
Those two numbers aren't actually contradictory, and that's the uncomfortable part. A pilot can produce a genuinely impressive 40% gain on a demo task. Controlled setting, engaged team, everyone paying close attention because they know they're being measured. Then it gets folded into normal operations, with normal attention spans and normal task variety, and six months later there's nothing left to show anyone who asks. The gain was real and it was local. It didn't generalize, and nobody built the measurement infrastructure to notice the gap between "worked in the pilot" and "still working now" before the renewal invoice showed up asking for another year of the same $200-a-seat commitment.
This is where hitechies' 2026 budget analysis makes an important, under-discussed point: the cost side of this equation is easy to track — seats, tokens, overages, all itemized on an invoice — while the value side stays soft indefinitely, because nobody committed to a hard number before the tool went live. Costs get audited by default. Value only gets audited by choice, and most orgs never make that choice.
This Is a Governance Failure, Not a Tooling Failure
None of this means AI coding assistants don't work. Refusing that framing matters. The lazy contrarian read — AI coding tools are a bubble, cancel your subscriptions — is just as unsupported by the data as the vendor claim it's reacting against. Developers hitting usage caps at a 30% rate are getting something out of these tools worth reaching for daily. The problem sits one layer up. It's in what companies chose to measure, which was mostly nothing, and what they chose to trust instead, which was mostly how it felt in the room.
This connects to something I wrote about last week on a different axis of the same failure: the psychological cost of AI-generated code that gets rewritten within two weeks and debugged by people who never built it. That piece was about what happens inside a developer when the struggle that used to build competence gets automated away. This one is about what happens inside an organization when the outcome that used to justify a purchase gets replaced by a feeling. They're the same failure wearing two different badges. One is a fatigue problem nobody measured, because there was no metric for whether a person still understood their own system. The other is an ROI problem nobody measured, because there was no metric for whether the work was actually faster — only a room full of people who agreed it seemed that way. Both failures share a root cause: a category that scaled faster than anyone's ability, or willingness, to check its own claims.
So Actually, the Category Isn't Broken. The Audit Is Missing.
The honest position isn't "AI coding tools deliver." It isn't "AI coding tools are overhyped," either. It's that an entire purchasing category, worth real budget line items across most of the software industry, is running without an audit trail. Companies wouldn't accept "the team feels like revenue is up" as a substitute for a revenue number. They're accepting the exact equivalent for engineering velocity, every renewal cycle, because the tool sits inside the black box of individual developer workflow — a box that's had poor visibility from outside long before AI showed up, which is exactly why it's been so easy to fill with a feeling instead of a fact.
The fix isn't a new dashboard vendor promising to solve this for another $200 a seat. It's the boring, unglamorous discipline of picking a real before-and-after metric. Cycle time. Defect escape rate. Time-to-first-review. Something that existed before the tool arrived, measured across a sample large enough that a genuine 19% slowdown wouldn't hide inside the noise of "everyone says it feels great." Most orgs haven't done this. It's slower and less flattering than citing the vendor's number, and until a team asks the question out loud, nobody has to know the answer might be uncomfortable.
The next time someone tells you their team is faster now, ask what the number was before. Not the mood before. The number. If nobody can produce one, you haven't found a team that's slower. You've found a team that never checked, and mistook the not-checking for good news.
The $200 invoice arrives whether or not anyone can prove what it bought. That's not a productivity story. It's a story about how easily an entire industry agreed to stop asking for receipts.