Data is “dark”
Stored but never used: no titles, descriptions, tags, quality ratings or privacy flags.
Gartner, IBM, Splunk
Data Discovery & AI Readiness
Privacy-first, AI-native metadata and data enrichment, at enterprise scale. Dtechtive turns fragmented, hidden data (whether open, internal or shared) into discoverable, trustworthy, compliant and usable datasets, over 100x faster than manual approaches, with all processing inside your own boundary.
Dtechtive built a strong partnership with the Data Platforms Team, helping improve metadata standards, support a major data migration, and enhance data discoverability while understanding the constraints of public sector clients.
Through the CivTech programme, the Scottish Government’s Digital Directorate and NatureScot collaborated with Dtechtive to create find.data.gov.scot, addressing the challenge of public sector data discoverability. This innovative solution not only offers enhanced metadata and intelligent search algorithms but also provides valuable insights into user behaviour. The CivTech process enabled a successful partnership with Dtechtive, resulting in a cutting-edge tool for accessing a vast array of Scottish public data.
Missing metadata, unstructured formats, no quality signals and compliance gaps prevent humans and AI agents from using data reliably, delaying data-driven decisions and increasing the risk of inaccuracy.
Stored but never used: no titles, descriptions, tags, quality ratings or privacy flags.
Gartner, IBM, Splunk
The global cost of storing and managing dark data that delivers no value.
DataStackHub, 2025
The average cost of poor data quality to a single organisation annually.
Gartner
The share of revenue firms lose to bad data across operations and decisions.
MIT Sloan
Without a semantic layer of context, AI hallucinates, retrieves the wrong sources, or stalls entirely.
Data consumers spend more than half their week searching, understanding and preparing data instead of using it.
Administrators and data owners cannot see what exists, what data teams search for, or what value their data delivers.
Most enrichment tools send data or AI inference to third-party clouds, creating compliance and security exposure.
When metadata, ownership and quality signals fall out of date, definitions stop matching how teams work. Trust and adoption decline, leaving another stale system of record instead of a reliable route to usable data.
Dtechtive transforms private/internal and public/web data, however fragmented, hidden or unstructured, into discoverable, trusted and AI-ready datasets. Months of manual effort become hours, and your data and AI inference never leave your environment.
01 – Enrich metadata
Self-hosted open-source models (e.g. Llama) or your own cloud AI generate rich metadata across spreadsheets, databases, maps and documents.
02 – Enrich data
Unstructured data becomes structured, analysis-ready and AI-ready, with every step running inside your boundary.
Dtechtive works alongside your existing infrastructure, turning data from cloud, inside your boundary, APIs and documents into trusted resources that power search and analysis.
Context-less and locked away across systems, in formats nobody can search.
Privacy-first, AI-native processing that slots in beside what you already run.
Discoverable, usable and governed across the whole data estate.
Natural-language semantic search across the full estate, with ranked results and filtering by quality, privacy, domain, owner and recency.
Generated overviews surface the most important context above search results so teams understand relevant data faster.
People and AI agents can ask natural-language questions and receive grounded answers without code or dashboards.
Most tools cover one half of the problem, in someone else’s cloud. Dtechtive covers both halves, inside yours.
We work across public/web, private/internal and shared data estates, using the same standards and approach either side of the firewall.
Solutions are fully deployed inside the client boundary: inside your boundary, or air-gapped. Data never leaves your environment, and AI inference runs there too.
We cover both, unlike other players who do one or the other. Context and the underlying data can both be enriched at scale.
An API-based modular tool that integrates with your existing catalogues, BI tools and warehouses. Open standards keep your metadata portable.
Over 100,000 datasets being enriched using privacy-first AI, from national open data to internal estates.
Case 1 · Open data
Scottish Government · Digital Directorate
open datasets · 70 data portals
Case 2 · Internal data
Scottish Government · Data Platforms Team
internal datasets · multiple sources
Case 3 · Contact data Confidential
Higher education & research
records · 15 years · 5+ sources
Privacy-by-design architecture. Your data and your AI inference never leave your boundary.
Privacy-by-design throughout, with data protection considerations built into every stage of enrichment and cataloguing.
Self-hosted or deployed in your own cloud tenancy (AWS, Azure, GCP). No third-party SaaS data transit.
Self-hosted Llama or bring-your-own cloud AI (Bedrock, Azure OpenAI, Vertex AI). No data sent to commercial LLM APIs.
Penetration testing and CyberEssentials certification underway for full enterprise assurance.
Pricing covers delivery stages rather than individual processing runs: enrichment runs take hours once deployed, depending on dataset volume and complexity; the pilot is 1–3 months, while integration, security review and rollout are separate months-long stages within the 6–24 month implementation.
01 · Evaluate
Experience in 30 minutes
Free
Demonstration using public data and selected AI features. Zero commitment.
Book your slot02 · Prove Most popular
1–3 months
£10k–£25k
Proof of concept on up to 1,000 datasets with specific formats and connectors, including privacy-first AI enrichment.
Scope a pilot03 · Deploy
6–24 months
Bespoke
Full-scale production deployment, unlimited datasets, advanced features tailored to your architecture.
Plan a rollout04 · Run
Ongoing
Bespoke
A base fee covering core platform access and standard support, with modular usage on top.
Get a quoteOn top of the base fee, usage scales across dataset volume (tiered by format complexity), AI processing credits (client cloud-native or client self-hosted), data connectors (structured, documents & BI, GIS & enterprise), user licences (Data Consumer, Data Steward/Owner, Admin), and a deployment fee based on hosting model. Optional add-ons: PII redaction, conversational search, synthetic data, and MCP server.
Straight answers on capability, sovereignty, standards and cost.
AI-native metadata and data enrichment uses AI models to automatically generate titles, descriptions, tags and quality scores for datasets, and to transform the underlying data into structured, analysis-ready formats. Dtechtive does both, where most tools cover only one, so an entire data estate becomes discoverable, trusted and AI-ready without manual effort.
Data sovereignty means your data never leaves your environment. AI sovereignty means AI inference also runs inside your boundary rather than being sent to a commercial LLM API. Dtechtive delivers both: it is deployed inside your boundary, using a self-hosted open-source model such as Llama, or your own cloud AI such as AWS Bedrock, Azure OpenAI or GCP Vertex AI.
Dtechtive is over 100x faster than manual enrichment. Manual metadata and data enrichment takes a skilled person at least one day per dataset; Dtechtive completes it in under a minute. Large estates that would take months to enrich by hand can be completed in hours once deployed, depending on dataset volume and complexity. Pilot, integration, security review and rollout are separate months-long stages.
No. Dtechtive uses an API-based modular architecture that plugs into existing catalogues, BI tools and data warehouses. There is no rip-and-replace and no vendor lock-in. Metadata is output in open standards including DCAT3, Croissant, Schema.org, Dublin Core and INSPIRE, so it stays portable.
Dtechtive works across spreadsheets and tabular data (CSV, XLSX), databases, documents (PDF, Word, HTML), geospatial and GIS formats (SHP, KML, GeoJSON, TIFF), and structured exchange formats (XML, JSON). It also extracts tables from unstructured documents and converts them into analysis-ready spreadsheets.
Yes. Dtechtive is built with privacy-by-design architecture, with data protection considerations built into every stage of enrichment and cataloguing. It detects sensitive information at row, field and document level, embeds privacy flags in metadata, and can redact or anonymise it. Because everything runs inside your own boundary, your data never transits third-party infrastructure.
Dtechtive starts with a free demo you can experience in 30 minutes. An enterprise pilot covering up to 1,000 datasets costs £10,000–£25,000 over 1–3 months. Full implementation and the annual licence are both priced bespoke: a base fee covers core platform access and support, with modular usage costs for dataset volume, AI processing, connectors, users and deployment.
Yes. Dtechtive works with public/open, private/internal and shared data. It powers find.data.gov.scot, covering 25,000+ open datasets across 70 Scottish public-sector data portals, and also enriches internal estates behind an organisation’s firewall.
See Dtechtive running on public data over a free demo, then scope a pilot on your own estate. No preparation needed.
Dtechtive is also on a mission to make open and commercial data more discoverable, trustworthy and AI-ready for everyone – not just inside the enterprise. Explore our public Web Data Search Engine.
Check out our Web Data Search Engine