Blog | Data Expo

Scalable AI Impact: From Trusted Data Sharing to AI-Ready Companies

Written by Data Expo | Aug 19, 2026, 12:57:05 PM

The bottleneck is rarely the model
Ask why an AI use case did not make it into production and the answer is seldom that the model underperformed. More often the project ran out of usable data, or ran into a question nobody could answer: are we permitted to use this dataset for this purpose, can we prove where it came from, what did we commit to when we received it from a partner.

That matters more as ambitions grow, because the data with the highest value tends to sit outside the organization. A manufacturer wanting to predict component failures needs data from suppliers. A hospital validating a diagnostic model needs cases from other hospitals. A logistics operator optimizing a corridor needs data from carriers, terminals and customers. Internal data is finite. Scaling AI impact means scaling access to data you do not own.

Why the usual route does not scale
The default answer is bilateral: two organizations negotiate a contract, agree what may be done with the data and build an integration. It works. It just does not compose. Every additional partner means another negotiation, another legal review, another integration. Effort grows with the number of relationships rather than with the value created, which explains a pattern many companies recognize: a proof of concept with one willing partner, followed by two years of not repeating it.

Scalable AI impact requires the next data relationship to cost far less than the first.

Trusted data sharing as shared infrastructure
This is what data spaces are for. A data space is an agreed set of rules and technical components that lets independent organizations exchange data while each keeps control of its own assets. Participants are onboarded once and verified against a common trust framework. Data stays where it is and is offered with a description of what it contains and what may be done with it. Those conditions then travel with the data instead of sitting in a contract that nobody consults while a system is running.

The shift is economic rather than technical. Because the terms are written once in a common form, admitting the next partner becomes closer to a repeatable process than a bespoke project. The rules are published openly as standards rather than owned by a vendor, and they are developed and maintained by the International Data Spaces Association (IDSA), a neutral non-profit association whose members include data users, technology providers and research organizations. IDSA sells nothing. Companies choose freely between open source and commercial implementations.

For an AI programme this matters in a specific way. Data arriving through a data space carries its own description of provenance and permitted use. What was previously reconstructed from email threads becomes an attribute of the asset itself.

What "AI-ready" actually requires
AI readiness is usually discussed as a technical property: clean data, sound pipelines, decent metadata, a platform that scales. Necessary, but incomplete once data crosses an organizational boundary.

At that point readiness has four dimensions, and interoperability is needed in all of them: technical (including syntactic), semantic, organizational and legal. Technical readiness means systems can exchange data. Semantic readiness means both sides mean the same thing by "delivery date". Organizational readiness means there are agreed roles and processes for onboarding, disputes and revocation. Legal readiness means the lawful basis, permitted purpose and downstream constraints are known and demonstrable.

An organization that is ready only in the technical sense is ready for pilots. Production AI on external data, with obligations under the General Data Protection Regulation (GDPR), the Data Act and the EU AI Act, needs the other three as well.

Making compliance keep pace with AI
Here a familiar trap opens up. Legal readiness is usually achieved by people: a review per use case, a data protection assessment per pipeline, an audit per year. That produces a defensible answer and reintroduces the scaling problem it was meant to solve. If every new AI use case costs another compliance project, governance becomes a tax on each initiative rather than a capability the organization owns.

Closing that gap is the subject of active European work. DataPACT, funded under Horizon Europe, sets out to embed compliance, ethics and sustainability into data and AI operations from the start instead of adding them afterwards. Its planned outputs include an AI Compliance Framework and an AI Compliance Toolbox intended to help organizations build compliance-aware data and AI pipelines, so that privacy and sustainability requirements are carried through the pipeline rather than assessed at the end. The project also works on open federated European data spaces as the setting for trusted cross-border exchange, with interoperability and traceability as explicit design goals, and is validating its results in sectors including healthcare, urban planning and marketing. IDSA is one of nineteen consortium partners and contributes to architecture development, requirements analysis, dissemination and impact creation.

Two things are worth taking from this even before the tooling matures. First, compliance evidence is becoming something a pipeline can produce as a by-product of running, which changes the unit economics of every subsequent use case. Second, sustainability is being treated as part of compliance rather than as a separate report, which is realistic given how much compute AI operations consume.

Where this is heading: participants that are not people
The direction of travel makes the scaling argument sharper. Software agents are starting to do what participants do in a data space: search a catalogue, negotiate terms, obtain data, act on the result.

IDSA's position paper Data Spaces and AI: Trustworthy Agentic Participation in Data Spaces examines what that requires. Its central finding is that no new and untested infrastructure is needed. The existing mechanisms suffice, applied to agents explicitly and consistently. Its central proposal is a Delegated Agent Participant ID, a Verifiable Credential linking an executing agent back to the organization that is legally liable for it, checked before a data transaction begins.

The relevance here is straightforward. Governance that depends on a human approving each exchange cannot operate at machine speed. Organizations that put automated, verifiable governance in place now are the ones for whom agentic participation will be an extension of normal operations rather than a new category of risk.

Three questions worth asking
A practical way to assess where an organization stands:

  • For each dataset feeding an AI system, can you state the lawful basis, the permitted purpose and the constraints attached by its source, without a research project?

  • Would onboarding your tenth data partner cost roughly what your second one did?

  • If an automated process used partner data eight months ago, can you show what it did and under whose authority?

Where the answers are uncomfortable, the constraint on AI impact is not model capability. It is readiness, and readiness can be built.

At Data Expo
Anil Turkmayali will present Scalable AI Impact: From Trusted Data Sharing to AI-Ready Companies at Data Expo, Lezingenzaal 3, on Thursday, September 10, 2026, from 16:00 to 16:30 at Jaarbeurs Utrecht. 

This blog post is a contribution from the International Data Spaces Association (IDSA), an international organization that invented the concept of data spaces and develops the global standard for them. For more information, visit www.internationaldataspaces.org or stop by the International Data Spaces Association booth at Data Expo.