I have given the impressive demo. I have also approved the budget after watching one. Both roles reward confidence before evidence, and I have been wrong in both.
Over the past 25 years I have started more than 30 companies. Twenty failed. That record taught me something that applies directly to corporate AI: the most dangerous moment in any investment is not when it visibly goes badly. It is when it looks like it is going well and nobody can prove it. That distance, between what an AI system reports and what it can be shown to have produced, is the AI validation gap.
Eighty-Nine To Six
Atlassian's Teamwork Lab surveyed 12,035 knowledge workers and 173 Fortune 1,000 executives for its 2026 State of Teams report. Eighty-nine percent of those executives said AI increases speed. Six percent said they were sure they had clear examples of organization wide AI ROI.
Eighty-nine to six. Now look at spend. DX, drawing on more than 500 organizations, found median quarterly AI spend rising from roughly $1,500 to nearly $44,000 in a year. In the technology sector it rose roughly 28 times. Over those same quarters, the innovation ratio, the share of engineering effort going to new capability rather than maintenance, stayed essentially flat.
| Signal | What Moved | What It Proves |
|---|---|---|
| Executive belief in speed | 89% say AI increases speed | Sentiment only |
| Evidence of ROI | 6% can cite clear examples | No proof |
| Median quarterly AI spend | ~$1,500 to ~$44,000 in a year | Cost confirmed |
| Technology sector spend | Roughly 28x increase | Cost confirmed |
| Innovation ratio | Essentially flat | Value unproven |
| Developer Experience Index | Fell from 67 to 65 over four quarters | Contradicts the dashboard |
Spend up twenty-eight times. New value creation up roughly one point. I am not arguing that AI does not work. It plainly does. I am arguing that a great deal of money is moving fast, and few people can produce evidence about where it landed.
Deployment Theater
Here is the pattern. An initiative gets approved on the strength of a demonstration. It gets built. It launches. There is an announcement, an internal celebration, maybe a press mention. Then attention moves on. Nobody was ever assigned to find out whether it worked.
"Every incentive is satisfied at launch, and none extends past it. The organization learns, correctly, that the announcement is the deliverable."
— Steve Taplin, on deployment theater
The important thing is that nobody is lying. The CEO needs an AI story for the board. Product needs one for the roadmap. Engineering needs the pressure to stop. That lesson is the real damage, because it makes the next initiative fail faster. Value producing AI exists, but from the announcement alone, real and staged look identical. The celebration carries no information.
↑ An Actual AI Strategy
- Named business metric with a baseline
- Holdout, phased rollout or matched comparison
- Cost measured after human review
- A tested stop mechanism with an owner
- Accountability that extends past launch
- Value proven before the next funding round
↓ An AI Announcement Strategy
- Funded on the strength of a demo
- Model accuracy and usage on the dashboard
- Verification time never counted
- A kill switch nobody has rehearsed
- Attention ends at the press mention
- Eighteen months past the point somebody knew
The Signal Read Backward
When usage of your AI system rises, that may be the failure signal. People use a system more when they trust it. They also use it more when they have to check its work, and they contact support more when it cannot resolve their problem. All of it shows up as engagement.
DX found the same structure inside engineering teams. Its Developer Experience Index fell from 67 to 65 over four quarters while developers' perceived rate of delivery stayed flat, even as measured output rose. Two instruments pointed at the same organizations disagreed for a year, and the one on the executive dashboard reported improvement.
Most AI dashboards measure model performance, usage and milestones. None of the three is a business metric, and all three can improve while value declines.
Why Capable Leaders Miss It
These are not naive organizations. They have boards, auditors and skeptical CFOs. So why does an initiative run 18 months past the point where somebody inside knew?
"The most common root cause of failure was the business leadership of the organization misunderstanding how to set the project on a pathway to success."
— RAND, interviews with 65 data scientists and engineers with 5+ years building AI and ML models
That matches what I hear. Across more than 200 recorded conversations with CTOs and founders on my podcast, the failures were rarely about model quality. They were about nobody defining what "working" meant, and nobody accountable for finding out. The signals that would tell you an initiative is failing are the ones your organization is built to suppress, because the person closest to the problem is the person whose performance depends on it succeeding.
Four Questions For This Week
1 — What Number Does This Move, And What Was It Before We Started?
If an initiative cannot name a business number and its baseline, it has failed the first test. Ask before funding, not after launch.
2 — What Would Have Happened Without It?
A result can improve after an AI implementation without improving because of it. Absent a phased rollout, a holdout or a matched comparison, what you have is correlation in a nicer suit.
3 — Who Can Turn It Off, And Have They Rehearsed It?
A stop mechanism nobody has tested is a diagram, not a control.
4 — What Does It Cost After Human Review?
A system that saves an employee four hours a week and requires three hours of checking did not create four hours of value. It created an impressive slide.
The Net Value Math Nobody Runs
Ask those in your next review. In my experience, the room goes quiet, and the silence is the finding.
The Shift
Adoption is finished as a differentiator. DX reports AI adoption above 90 percent in the engineering sector, so there is no longer a control group to compare against. The next phase will be won by whoever proves value fastest. Your competitor can rent the same models you rent within a quarter. What does not commoditize is knowing what is actually working.
I am not here to slow anyone down. I am here to stop companies from scaling fiction. Carry one question into your next board meeting: "Can we prove this works in the real world?"