The Data Catalog: Why Nobody Can Find the Data They Already Have
Here's a question that quietly costs insurers a fortune: when someone in your organisation needs data, how do they find out whether it already exists?
At most insurers, the answer is "they ask around." They message a few people, someone half-remembers a table, someone else built something similar last year, and after a few days of archaeology they either find it or — more often — give up and rebuild it. The data existed the whole time. Nobody could find it.
This is the discoverability problem, and a data catalog is the answer nobody prioritises until they realise how much they're paying to not have one.
The cost of "ask around"
Duplicated work. A team spends weeks building a dataset that already existed in another team. Multiply across an organisation and a meaningful fraction of data effort is spent recreating things that were already there.
Duplicated data. Worse than duplicated effort — now the same data exists in multiple places, maintained separately, drifting apart, producing inconsistent answers. (And, as it happens, inflating your cloud bill.) The four teams each keeping their own copy of policy data, because none could find or trust a shared one, is this problem in physical form.
Slow everything. Every project that needs data starts with a discovery phase that shouldn't exist. Analysts and scientists spend more time finding data than analysing it — a widely-cited and entirely avoidable tax.
Knowledge that walks out. Where the data is and what it means lives in people's heads. When they leave, the map leaves with them — the same tribal-knowledge risk as legacy systems, applied to the data estate itself.
Wrong decisions from the wrong data. Unable to find the authoritative source, people use whatever they can find — an old extract, a personal copy, a table that looks right. Decisions get made on data that was never meant for that purpose, and nobody knows.
What a data catalog actually is
A searchable inventory of your data. For each dataset: what it is, where it lives, what it means, who owns it, how fresh and how good it is, where it comes from, and how it may be used. A search engine for your own data estate.
The value is simple: someone who needs data searches, finds what exists, understands whether it fits, and uses it — instead of asking around, giving up, and rebuilding.
Why insurers especially need one
Sprawling, decades-old estates. Data across many systems accumulated over decades, with layers of extracts and copies. The more sprawling the estate, the more impossible discovery-by-asking becomes.
Regulatory demands. Increasingly you must know what data you hold, where it is, and how it's used — for privacy, for AI governance, for reporting. A catalog is part of the substrate those requirements assume.
The AI dependency. AI projects need to find the right data. A catalog is how teams discover what's available to build on, instead of each one rediscovering the estate from scratch.
Institutional memory loss. As experienced staff retire, undocumented knowledge of the data estate is being lost. A catalog captures it before it walks out the door.
Why most catalog projects fail
Worth being honest, because plenty of insurers have bought a catalog and gotten no value.
1. Empty catalog syndrome. The tool gets deployed, and it's empty — or auto-populated with technical metadata (table and column names) that's useless without meaning. A catalog that lists TBL_PLCY_MSTR.FLD_47 with no explanation helps no one. It needs the human context: what this is, what it means, whether to trust it.
2. It goes stale immediately. A catalog populated once and never maintained is wrong within months and quickly ignored. It has to stay current — through automation where possible, and ownership everywhere else.
3. No ownership. The recurring theme. Catalog entries need owners who keep them accurate. Without the underlying ownership model, the catalog decays into an unreliable list nobody trusts.
4. Bought as a tool, not adopted as a practice. A catalog is only valuable if people actually use it to discover data and actually maintain it. That's a behaviour change, not a software install — and behaviour change is the part that gets skipped.
How to make it work
Start with the data people actually look for. Don't catalog everything. Catalog the most-used, most-valuable datasets first — the ones people currently ask around for. Immediate value, visible payoff.
Include meaning, not just structure. The context — what it means, whether to trust it, how to use it — is the point. Technical metadata alone is the empty-catalog trap.
Connect it to ownership, quality, and lineage. A catalog entry with an owner, a quality score, and lineage is genuinely useful. Isolated, it's a phone book. These aren't separate projects; they're one data-governance capability.
Automate what you can, assign owners for the rest. Automated metadata keeps the basics current; human-curated meaning and ownership keep it trustworthy.
Make discovery the default. The catalog only pays off when "search the catalog first" becomes the reflex before anyone builds new data. That's the behaviour that turns it from shelfware into savings.
The point
Insurers spend enormous effort producing data and comparatively nothing making it findable — then pay for that omission every day in duplicated work, duplicated data, slow projects, and decisions made on whatever could be found.
A data catalog isn't glamorous. It's a search engine for data you already have. But in a sprawling, decades-old insurance data estate, being able to find what exists — and know whether to trust it — is worth more than most of the tools that get funded ahead of it. You already own the data. The catalog is how your people stop rebuilding it.
We build data catalogs, ownership models, quality scoring, and lineage as one connected governance foundation. More at IntelliBooks.
Comments
Post a Comment