IBM spent $4 billion learning that a cancer AI couldn't read a real chart.
IBM's Watson for Oncology was trained using hypothetical cases created by a small group of specialists.
Then reality showed up.
Real patient records were messy, contradictory, and full of details the system hadn't learned to handle. When doctors at other hospitals tested it, some recommendations were described as unsafe and incorrect.
The project was eventually shut down.
Zillow Offers had a different problem.
Its home-buying algorithm was trained during a relatively stable housing market. When prices became unpredictable in 2021, the model continued making purchases that later had to be sold at losses—costing Zillow roughly $500 million.
Then there was Builder.ai.
The company raised money at a $1.2 billion valuation around the promise of AI-powered software development. Reports later indicated that much of the work depended on human engineers. The company filed for bankruptcy in 2025.
Different companies.
Different industries.
Different technologies.
Same underlying mistake.
They started with AI instead of the problem.
This is the pattern that appears again and again.
Someone says:
“We need AI”
Then the organisation starts searching for somewhere to use it.
That's backwards.
The source material cites MIT's 2025 NANDA research covering 300 enterprise AI deployments, where 95% of generative AI pilots showed zero measurable P&L impact.
The important question isn't simply whether the model works.
It's whether the model works inside the workflow where the business actually operates.
A hospital doesn't have a clean dataset.
A bank's fraud team doesn't follow a perfectly predictable process.
A customer-support team doesn't handle every ticket the same way.
Every real workflow contains exceptions, judgment calls, politics, shortcuts, and tribal knowledge.
A generic AI model doesn't automatically understand any of that.
And that's why the successful projects looked different.
They weren't trying to "transform the business with AI."
They focused on one narrow task.
One measurable outcome.
And one business owner accountable for making it work.
Before funding your next AI pilot, ask 3 questions.
1. What exact task are we changing?
You should be able to describe it in one sentence.
"Improve customer service" is too vague.
"Reduce average ticket resolution time from 14 minutes to 6" is measurable.
If you can't define the task, you probably aren't ready to build the AI.
2. Whose job gets easier?
Identify the person who actually uses the system.
Then involve them before development—not after.
The people closest to the workflow usually know the exceptions that never appear in the requirements document.
3. What happens when the model is wrong?
This question is often ignored.
Does the system fail safely to a human?
Or does it quietly make the wrong decision in production?
AI doesn't need to be perfect.
But the workflow around it needs to be designed for failure.
And set a kill date.
One of the easiest ways for an AI project to become expensive is to never officially fail.
The pilot keeps getting extended.
Another experiment.
Another proof of concept.
Another meeting.
Eventually, the organization has spent months—or years—in pilot purgatory without making a production decision.
A simple solution:
Set a go/kill review at week eight.
If your chosen metric hasn't moved by then, stop.
Learn.
Change the hypothesis.
Or kill the project.
Don't let "we're still testing" become the permanent status.
The real AI advantage isn't the model.
It's knowing where to apply it.
The companies that win with AI won't necessarily be the ones with the biggest models or the largest AI budgets.
They'll be the ones that can answer three questions clearly:
What specific work are we changing?
Who owns the outcome?
When will we know whether it worked?
Because the problem isn't that AI doesn't work.
The problem is funding AI projects before knowing what success actually means.
The failure rate isn't just a technology story.
It's an organizational discipline story.
One question to take into your next AI meeting:
What specific work gets easier, for whom, starting when?
If nobody can answer it clearly, you may not have an AI project yet.
You may just have a demo looking for a problem.
The question for you
What assumption inside your organization's AI roadmap have you never actually tested?
If this changed how you'll approach your next AI investment, forward it to the person who signs off on the budget.
Until next week,
Keep questioning the obvious.
