AI & Knowledge

Data is not ready for AI: how to tell if it is (and what to do)

3 min read
Data is not ready for AI: how to tell if it is (and what to do)

The most advanced model in the world produces mediocre results if the data and knowledge you give it are incomplete, duplicated, or obsolete. Data readiness is often the real bottleneck—not the choice of model.

What “ready data” means in practice

It does not mean “having a perfect data lake”. It means, for the scope you want to automate or assist:

  • critical information exists and is findable;
  • there is a recognizable current version;
  • access is coherent with risk;
  • sources are citable and maintainable;
  • quality is sufficient for the reliability level you need.

Gartner and other 2025–2026 analyses continue to point to data quality and availability among the main causes of AI project abandonment. Without this base, the pilot looks brilliant and production does not.

Checklist of data and knowledge readiness for AI

An operational six-point checklist

  1. Findability
    Can people (and systems) find the right document or data in acceptable time?

  2. Current version
    Is there a clear golden source, or do three versions of the same file coexist?

  3. Completeness for the use case
    Is all information needed for the process present, or are critical pieces missing?

  4. Accessibility and permissions
    Who should see what? Are rights aligned with risk?

  5. Update and ownership
    Who is responsible for keeping that knowledge current?

  6. Citability and verifiability
    Can AI (or a person) indicate where the information comes from?

If more than two points are weak on a process, the AI project will start uphill.

Why SMEs suffer more (and how to limit the damage)

In SMEs, critical knowledge often lives in:

  • messy shared folders;
  • email and chat;
  • people’s heads;
  • unofficial working Excel files.

You do not need to fix everything. You need to fix the scope of the first use case.

The method that works:

  1. Choose a process.
  2. List the 10–20 documents or sources that feed it.
  3. Put order only on those (versions, owner, access).
  4. Only then connect AI.

This is also the through-line of the Zendata series: usable knowledge before the tool.

Signals that data is not ready

  • AI answers are plausible but “do not match” operational reality.
  • People keep asking colleagues “which is the good version”.
  • The pilot only works with hand-selected data.
  • You cannot cite sources reliably.

FAQ

Do you need months of data cleaning?
No. Start from the use-case scope. Prove value and expand.

Are structured data (ERP/CRM) enough?
Rarely. Decision and procedural knowledge often lives in unstructured documents.

Can AI help put things in order?
Yes, it can support finding duplicates, classification, and suggestions. Human responsibility for versions and decisions remains necessary.

When is data “ready enough”?
When, on that process, answers become citable, verifiable, and useful in daily work—not only in a demo.

Sources

Dig deeper in the series

If you want a quick readiness assessment on a specific process (documents, access, versions, gaps), we start there. Write to info@zendata.it or visit zendata.it.

Pietro Ciattaglia, CEO of Zendata AI, Rome