Case study · Internal data
Enriching 75,000+ internal datasets without data leaving the boundary
An internal data estate at national scale, spread across multiple sources, with sensitive content and no consistent metadata. Enrichment had to happen inside the client's own environment.
- Client
- Scottish Government Data Platforms Team
- Sector
- Public sector · Internal data
- Deployment
- Client cloud tenancy, sovereign
The challenge
Where the open data problem was about visibility across organisations, the internal problem was about scale and control within one.
The estate ran to tens of thousands of datasets drawn from multiple sources, accumulated over years and in wildly varying condition. Much of it had no usable metadata at all: no descriptions, no tags, no quality signals, no indication of whether a file contained personal information. At that volume, manual cataloguing is not slow. It is arithmetically impossible. A skilled person takes roughly a day per dataset; 75,000 datasets is several lifetimes of work.
Two constraints made the usual answers unavailable. First, sensitive content meant the data could not be sent to a commercial AI service for processing. Second, a major data migration was underway in parallel, so the solution had to work with a moving target rather than a settled estate.
Our approach
- Sovereign by default. The tool was deployed inside the client's own cloud tenancy. Data never leaves the environment, and AI inference runs there too, using a self-hosted open-source model or the client's own cloud AI service. No third-party data transit, no exposure to commercial model training.
- Enrich both metadata and data. Generating descriptions is only half the job. Formats were harmonised and structures normalised so that downstream systems could actually consume what the catalogue described.
- Detect sensitive information early. Personal information is identified at row, field and document level, with privacy flags embedded in the metadata itself so they travel with the asset.
- Standardise for interoperability. Output was aligned to DCAT3, so enriched metadata is portable across government systems rather than locked into one tool.
- Keep humans in the loop. The model drafts; data owners approve, edit or reject. That preserves the organisational nuance an automated system never quite learns.
The solution
The Metadata Enrichment Tool applies AI-native enrichment across the estate at a pace no manual process could match, generating titles, descriptions and tags, scoring quality across completeness, validity, consistency and uniqueness, and flagging sensitive information, all within the client's boundary.
Because it operates as an overlay rather than a replacement, it required no migration of the underlying data and no change to the systems already in place. It runs alongside the existing estate rather than asking the estate to reorganise around it.
Results
The engagement demonstrates something the market often treats as a trade-off: enrichment at scale and full data sovereignty are not mutually exclusive.
- An estate that can be searched. Datasets that were previously invisible to the people who needed them are now catalogued, described and findable.
- Compliance posture strengthened. Sensitive information is identified systematically rather than discovered incidentally, with privacy flags carried in the metadata.
- Interoperability by design. DCAT3 alignment means the metadata can move between government systems without rework.
- A foundation for AI. The enriched catalogue is the semantic layer that downstream AI and analytics need in order to produce reliable answers.
Data Platforms TeamThe Scottish GovernmentDtechtive built a strong partnership with the Data Platforms Team, helping improve metadata standards, support a major data migration, and enhance data discoverability while understanding the constraints of public sector clients.