It is precisely when you scale generative AI that issues with data quality and data governance start to surface – issues that may still seem manageable during a pilot. Outdated documents, inconsistent definitions and unclear ownership have a direct impact on the reliability of AI-generated answers.
This also affects the business case. If employees get an answer quickly but then have to spend a significant amount of time checking whether it is correct, much of the expected time saving disappears.
In this article, we explain why AI exposes existing data problems and explore three key data questions you need to address to move AI reliably from pilot to practice.
TL;DR
AI is only as reliable as the information it relies on. As you scale, issues with outdated information, inconsistent definitions and unclear ownership quickly become apparent. Three things are therefore essential: clear ownership, traceable information and ongoing attention to data quality.
Why data problems become visible when you scale AI
Setting up an initial AI application is relatively straightforward these days. Organisations are experimenting with Copilot, connecting language models to internal documents or developing AI assistants for specific processes.
This can work well within the controlled environment of a pilot. The real test begins when you add more users, data sources and processes.
You may discover, for example, that marketing and finance use different definitions of an ‘active customer’. Or that multiple versions of the same document are stored in SharePoint. Customers or assets may be recorded differently across systems, while for some data sources it may not even be clear who is responsible for them.
People can often interpret these inconsistencies. Based on experience, they know which document is current or which definition their department uses. AI does not automatically have that context.
As a result, an AI project can inadvertently become a stress test for your data foundation.
Why a convincing AI answer can still be wrong
The principle of garbage in, garbage out is nothing new. Generative AI does, however, make its consequences more visible.
An AI system can turn outdated or incorrect information into a convincing answer, without making it immediately obvious to the user that the underlying source is unreliable. This becomes particularly important when employees use AI for customer advice, maintenance, analysis or decision-making.
Research also shows that poor data quality is a barrier to scaling generative AI. Gartner reported in early 2026 that, by the end of 2025, at least 50% of generative AI projects had been abandoned after the proof-of-concept stage. Gartner cited poor data quality, inadequate risk controls, escalating costs and unclear business value among the reasons.
Unstructured data also deserves attention. Harvard Business Review highlights the importance of proprietary company information – such as emails, documents, contracts and SharePoint files – in determining the value that generative AI can deliver.
Data quality, therefore, is not just about databases. For documents, too, it needs to be clear which version is current, who is responsible for the content and which information AI is allowed to use.
Three data questions to answer before scaling AI
You do not need to have your entire data landscape in perfect order before you start using AI. But if you want to scale, you need to know which data is critical and how you will manage its quality. Three questions can help.
1. Who is responsible for the data?
For critical data sources, it should be clear who is responsible for their quality and accuracy.
If an AI assistant uses an outdated procedure, who decides which version is valid? And who corrects the source? Without clear ownership, errors persist and the reliability of the AI application remains vulnerable.
2. Can you trace where information comes from?
For business-critical applications, you need to be able to determine what an AI-generated answer is based on.
This requires clear definitions, metadata and visibility into the origin of information. It enables employees to better assess whether an answer is based on the correct and most up-to-date source.
Traceability is not just a matter of quality, either. Since August 2026, certain AI systems have been subject to requirements under the European AI Act relating to areas including transparency and documentation. The exact requirements depend on the application and your organisation’s role. Good data governance helps you maintain visibility into the information an AI application uses and how that information is managed.
3. How do you maintain data quality?
A one-off clean-up before an AI application goes live is not enough. Data is constantly changing.
That is why you need to determine which sources are critical, what quality standards they need to meet and who takes action when quality deteriorates. This makes data quality part of ongoing management rather than a one-off exercise.
Practical example: How data quality affects an AI assistant in practice
Imagine a grid operator introduces an AI assistant to help maintenance engineers find technical information more quickly. The pilot works well: engineers ask questions in natural language and find the right instructions faster.
As the solution is rolled out more widely, however, inconsistencies start to appear in the answers. The AI model is not the cause. The same maintenance procedure is stored in several locations, and not every version is up to date. Asset data is also recorded differently across systems.
A better AI model will not automatically solve these problems. What is needed first is clear ownership, consistent definitions, effective version control and ongoing monitoring of data quality.
Reliable data is a key part of the AI business case
An AI pilot can work well from a technical perspective and still deliver little value. If employees receive a quick answer but then have to spend considerable time checking it, much of the time saving disappears.
A reliable data foundation reduces the need for checking and correcting information, helps manage risks and makes further scaling possible. Reporting, dashboards and other data-driven processes benefit from this too.
With our AI Opportunity Scan, we identify which AI use cases offer the greatest business value and what is needed in terms of data, technology and governance to implement them successfully.
The fact that AI exposes existing data problems mainly shows where improvements are needed. So do not just ask what the technology can do. Also ask: is our data good enough for us to trust the outcome?
That is a key factor in determining whether an AI pilot develops into an application that delivers genuine value.
![]()
This blog post is a contribution from Digital Power, your data and AI partner. Digital Power helps you gain control over your data and brings AI into practice. They build scalable, secure, and future-proof solutions. For more information, visit www.digital-power.com or visit Digital Power during Data Expo.
This is an article by Guus van Loon (Lead Consultant Strategie & Transformatie – Digital Power)
With more than 10 years of experience in data, strategy and innovation, Guus helps organisations turn complex ambitions into sustainable business value. As Head of Research & Development at a market research scale-up, he worked on strategic challenges for international clients. He is a valuable bridge between technical specialists and business stakeholders, as well as a trusted adviser to senior management.