Data and AI sovereignty #
Data sovereignty means an organisation's data never leaves its own environment. AI sovereignty extends the same principle to inference: the AI model runs inside that boundary too, rather than sending data to a third-party commercial API. Together they allow an organisation to adopt AI without relinquishing custody of its data, which is why it matters most in government, healthcare, defence and financial services.
AI-native enrichment #
AI-native enrichment is the use of AI models to automatically generate descriptive metadata, score data quality, detect sensitive information and restructure raw files, rather than performing these tasks manually or with rule-based scripts. The distinction from "AI-enabled" is architectural: an AI-native tool is built around model inference as the primary mechanism, not as a feature added to an existing catalogue.
Agent-ready data #
Agent-ready data is data an autonomous AI agent can find, interpret and use without human intervention. It requires machine-readable metadata, explicit provenance, quality signals and clear usage rights. Publishing data openly is not sufficient: without this context an agent cannot judge whether a dataset is relevant, current or trustworthy, so it either ignores it or uses it badly.
DCAT3 #
DCAT3 is the Data Catalog Vocabulary, version 3, a W3C standard for describing datasets and data services so they can be shared between catalogues and organisations. It is the basis of the UK Government Metadata Exchange Model and is widely used across European public-sector data infrastructure, which makes it the practical default for cross-government data sharing.
Read the W3C specification →
Croissant #
Croissant is an open metadata format from MLCommons that describes machine learning datasets, combining metadata, resource descriptions, data structure and default ML semantics in a single file. Built as an extension of Schema.org, it makes datasets portable across ML frameworks such as PyTorch, TensorFlow and JAX, and discoverable beyond the repository hosting them.
Read the MLCommons specification →
Dark data #
Dark data is data an organisation collects and stores but never uses, typically because it lacks the metadata needed to find, understand or trust it. Industry estimates from Gartner, IBM and Splunk place dark data at between 55% and 80% of data holdings, representing both a wasted asset and an ongoing storage, governance and compliance cost.