Fragmented data to smart decisions

AI-assisted metadata enhancement

Raw data is a pile of bricks without a blueprint. Metadata is the blueprint, and AI can now draft it at a scale no manual process can match.

Dr Gautham Krishnadas ·19 January 2025 ·6 min read

In this age of data and AI, it is easy to assume that raw numbers and variables tell the full story. But anyone who has spent time navigating datasets, whether spreadsheets, maps or databases, knows the reality: raw data is often like a pile of bricks without a blueprint. To make sense of it, we need something that provides context and structure. This is where metadata comes in.

Metadata is the key to realising the true value of data. In its simplest form, metadata is information about data. It tells us what the data is, how it is organised, and who created it. Think of it like the table of contents of a book: it helps you understand what is inside before you even start reading.

Dataset. A collection of related data, often organised in tables or files, that can be used for analysis to derive insights – spreadsheets, maps, documents or databases.

Data asset. A broader term that includes datasets, but also extends to other data resources such as metadata, data models, visualisations or code.

Why metadata is essential, but often overlooked

Here is the thing: without good metadata, data can be a maze. It is like walking into a library where every book is just a stack of pages, with no labels and no author names. How do you find what you are looking for? Well-structured metadata makes it easy to search, understand and use data.

And it is not just about making data easier to find. Metadata provides critical context. A good description, clear tags and reliable data provider information can tell you whether a spreadsheet is worth your time before you even open it. The right metadata answers your questions in advance.

Despite its importance, metadata is often the last thing people think about. In fact, our research found that fewer than 30% of open data portals follow good metadata practices. That is a big problem. Without the right context, data becomes a jumble of numbers, making it harder to understand, use or collaborate on. Worse, poorly structured metadata leads to wasted resources and missed opportunities. Much of the data goes underutilised in a world that is not only meant to be data-driven, but is also on the brink of AI disruption.

The role of metadata standards

Metadata standards define how metadata is structured, labelled and formatted. For a fuller treatment, see Metadata standards today and beyond.

Schema.org is a standard that helps describe web data in a way search engines and systems can easily understand. It acts as a universal language for online data.

DCAT improves data discovery and sharing by providing a common model for describing datasets, metadata and their relationships in a catalogue.

As a practical example, Google Dataset Search – the world's largest web data search engine – can identify datasets on web pages that use either Schema.org or DCAT metadata. Going back to our research, fewer than 30% of websites that publish open datasets have adopted either standard. That highlights a significant gap in web data discovery through mainstream search engines. Adoption of metadata standards within organisations is not expected to be any better.

AI-assisted metadata enhancement

So what happens when metadata is sparse or absent, and adopting metadata standards seems like a distant dream?

Dtechtive's AI-assisted metadata enhancement scans the content and structure of your data, whether spreadsheets, maps, documents or databases, to generate metadata that follows standards such as DCAT and Schema.org. It assesses the quality of both the data and its metadata against internationally recognised frameworks, and identifies sensitive information such as email addresses and phone numbers. Most importantly, humans are kept in the loop to accept, reject or edit the AI-generated metadata.

The result? Data that is easier to find, understand, trust and use.

How this benefits you

Web data discovery

The metadata enhancement tool powers Dtechtive's web data discovery, bringing datasets hidden from mainstream search engines into the light. There are millions of such hidden datasets, and we are on a mission to make them discoverable and usable. You can explore this at dtechtive.com/search.

Internal data discovery

Dtechtive also uses the same capability to make data assets inside organisations more findable, understandable, trustworthy and usable. The real kicker here is the use of privacy-first AI that can be self-hosted, giving organisations full control over their data. This ensures that no internal data is exposed to commercial AI, protecting privacy and ensuring security.

Wrapping it up

To sum it all up: good metadata is not just a nice-to-have, but a need-to-have for ensuring data is discoverable and usable. By combining AI-generated and human-curated metadata, Dtechtive is realising the potential of underutilised web and internal data, driving data-driven innovation across industries.

Next step

See it working on your data.

A free demo shows Dtechtive enriching real data in 30 minutes, and takes no preparation from your side.

Book a free demo