Documents hold a company’s true operational knowledge. Management systems record transactions; contracts, procedures, meeting notes, and important emails explain rules, exceptions, and responsibilities. If AI cannot access that information in an ordered way, it works on half the story.
What ERP and CRM record, and what documents explain
ERP and CRM are essential, but they mainly tell facts: a customer bought, an invoice was paid, a ticket was closed. On their own, they do not explain the “why” of a decision.
In documents you find instead:
- contractual constraints and special conditions;
- operating procedures and their exceptions;
- responsibilities and approvals;
- the context of choices made in meetings or by email.
Without this layer, even a well-written answer stays fragile: the context that matters in daily decisions is missing.
Why many Italian companies start with AI without the document base
In practice, critical knowledge is often:
- scattered across shared folders, email, and personal drives;
- hard to find and hard to update;
- full of obsolete versions and duplicates.
According to Istat, Enterprises and ICT 2025, only 15.7% of Italian SMEs with at least 10 employees use artificial intelligence tools, versus 53.1% of large enterprises. The gap has many organizational and skills causes: the quality and accessibility of document knowledge are a relevant component, not the only one.
OECD analyses on GenAI in SMEs also highlight barriers linked to data, processes, and internal capability, not only to model availability.
What happens when AI does not “see” the right documents
A powerful language model without ordered access to company knowledge tends to produce incomplete answers, based on old versions or hard to verify.
Here it helps to separate three levels:
- corpus quality: up-to-date, non-contradictory documents with a clear current version;
- retrieval: the ability to recover the right fragments;
- generation: the model’s ability to use that context without inventing.
The original Retrieval-Augmented Generation (RAG) paper shows exactly this mechanism: retrievable external memory can improve factuality and provenance, but only if the retrieved sources are relevant. Later studies on RAG robustness at the query level remind us that the system remains sensitive to how the question is asked and to what is actually found.
In short: order and accessibility of the corpus do not eliminate every error, but they reduce an important piece of fragility.
How to make a document perimeter usable
At Zendata we start from documents not because structured data does not matter, but because that is where the knowledge that makes the difference in daily decisions lives.
On a delimited perimeter, the work is to make that knowledge:
- Findable: with minimal metadata (owner, date, current version);
- Readable: for people and for systems;
- Accessible: with clear controls on who can see what;
- Up to date: reducing duplicates and obsolete versions.
Only then does AI become a natural way to access knowledge, not a fragile experiment.
How to start today, without remaking the whole company
You do not need a complete document revolution. Three steps are enough:
- choose a repetitive, frequent process (internal requests, customer care, reporting);
- map where the documents that feed it live;
- make that perimeter readable, up to date, and accessible.
On a scoped pilot, the first improvements in accuracy and speed can typically appear within a few weeks; this is not a universal guarantee and depends on perimeter, source quality, and adoption.
AI accelerates only if there is something solid to accelerate.
FAQ
Can unstructured documents feed reliable AI?
They can contribute substantially if organized with metadata, versioning, and access controls. RAG helps retrieve relevant fragments and cite sources, but it does not by itself guarantee correct answers if the corpus or retrieval is weak.
Do you need an expensive document system?
No. You can start from SharePoint, company drives, or existing folders, adding order, ownership, and accessibility progressively.
What is the biggest risk?
Leaving AI to work on half the story: plausible but fragile answers that are hard to defend against operational reality.
Why does Zendata start from documents?
Because that is often where rules, exceptions, and responsibilities live that transactional data alone does not explain.
Sources
- Istat: Enterprises and ICT, 2025: AI adoption in SMEs and large enterprises.
- OECD: Generative AI and the SME Workforce: GenAI barriers and opportunities in SMEs.
- Lewis et al.: Retrieval-Augmented Generation (arXiv:2005.11401): RAG foundations, external memory, and provenance.
- Investigating the Robustness of RAG at the Query Level (arXiv:2507.06956): retrieval sensitivity to query quality.
Explore the series
- When AI gets it wrong: the problem is often context
- Not which AI to choose: which knowledge to make usable
- Where to start with AI in the company
If you want to understand whether a process in your company has a document base clear enough for AI, we can start from a targeted assessment of that perimeter: sources, versions, access, and gaps. Write to us at info@zendata.it or visit zendata.it.
Pietro Ciattaglia, CEO of Zendata AI, Rome