Data Readiness for AI: Why Going Upstream Beats Cleaning Up


Two fishermen are sitting beside a river.

A basket floats past. There is a baby inside it. They jump in and pull it out. Then another basket comes. And another.

Thirty minutes later they are both standing in the water, bent over, out of breath, still pulling.

Then one of them starts walking away. Upstream.

The other one shouts after him. Where are you going?

He says: I am going to find out who is putting them in.

That story is usually credited to the sociologist Irving Zola. John McKinlay used it in 1975 to describe healthcare, and people have retold it ever since. I use it to describe data.

Because right now, most leaders I meet are standing in the river.

You are already in the water

Think about your last month at work.

You checked an AI output because you were not sure it was right. You corrected a report. You sat in a meeting where two departments arrived with two different revenue numbers, and nobody could say which one was correct.

That is not an AI problem. That is a data problem wearing an AI costume.

And it costs you twice. Once for the AI licence. Once for the people checking whether the AI is telling the truth.

Gartner puts the average cost of poor data quality at $12.9 million a year per organisation. That figure is from 2020 and it covers every industry, so treat it as a signal rather than your number. The point stands. This was expensive before AI arrived. AI just makes it faster.

What AI actually does to bad data

Here is a test you can run in five minutes.

Pull a list of customer names out of your CRM. A real list. Not a tidy sample somebody prepared for a demo.

You will probably find something like this:

  • J. Smith
  • John Smith
  • Jon Smith
  • SMITH, JOHN

Now ask an AI tool: how many customers do we have?

It will answer straight away. Clean formatting. Confident tone. Maybe a nice chart to go with it.

And it will be wrong.

The big idea

AI does not fix your data. It exposes it.

Old software failed loudly. Bad input, error message, red box, everything stops. You knew there was a problem.

AI fails quietly. Bad input, confident answer, tidy formatting. You do not know there was a problem.

It is like hiring an accountant who is sometimes wrong but never unsure.

That is not a technology problem. It is a trust problem, and it lands on your desk.

Data readiness for AI: not all mess is equal

This is where a lot of advice is now out of date.

The old rule was simple: clean everything first. That rule was written for older systems that needed neat rows and columns. Today’s AI tools read messy documents, emails and transcripts reasonably well.

So the question is no longer “is our data clean.” The question is “which kind of mess do we have.”

Type of mess Does it still break AI?
The same customer stored four different ways Yes
Information nobody ever wrote down Yes
No agreed definition of what a customer is Yes
Messy but accurate PDFs, emails and notes No, this has changed
Data spread across several systems No, manageable

The top three are upstream problems. They happen at the moment data is created, or at the moment it fails to get created at all.

The bottom two are downstream problems. Tools handle those now.

Most companies are spending their money on the bottom two.

The trap to avoid

At some point in this conversation, somebody will say the answer is a single source of truth for the whole business.

Ask them one question. How long will that take?

The honest answer is years. And while you wait, nothing happens. “We are not data ready” quietly becomes the reason to do nothing at all.

You do not need one source of truth for everything. You need one process where the data is right.

What leaders should do

  1. Pick one process. Not the company. One. Sales handover, invoicing, customer complaints.
  2. Ask where the data is born. Who types it in? When? Into what box?
  3. Fix the moment of creation. A dropdown instead of a free text field. A required field. Two minutes at the end of a call to write down what was agreed.
  4. Put one name on it. A person, not a department. Departments do not own things.
  5. Then point AI at it. In that order.

Four of those five steps have nothing to do with technology.

Bad data is not a technology problem. It is a habit problem. And habits belong to leaders.

Try this this week

Ask three people in your business the same question, separately. Something basic. How many customers do we have?

Do not ask them in the same room. Do not tell them why you are asking.

Then compare the three answers.

If they match, you are ahead of most companies. If they do not, you have just found your river.

The takeaway

AI will not clean up your data. It will tell everybody what is in it, faster, and with a straight face.

So stop pulling baskets out of the water.

Start walking upstream.


Ready to find your upstream?

I run hands-on AI workshops for leadership teams. We take one real process from your business, trace the data back to where it is created, and fix it there. No theory. No five year data programme.

Book a workshop

Share This Story, Choose Your Platform!