Third-Party Data: Buying Signals You Already Own

The third-party data market for insurers is enormous and growing. Property attributes, company financials, credit signals, vehicle histories, geospatial layers, behavioural scores — there's a vendor for everything, and most carriers are buying more of it every year.

A significant portion of that spend is wasted. Not because the data is bad, but because of two failures that repeat with striking consistency: buying signals you already own, and buying data you can't join.

Failure one: buying what you already have

An insurer buys a behavioural risk score. Reasonable purchase. But a good share of what drives that score is derived from things already sitting in their own systems: claims frequency, payment behaviour, policy tenure, contact-centre interaction history, cancellation history.

They're paying a vendor for a modelled proxy of their own customer's behaviour — because their own version of that behaviour is scattered across four systems with no customer key, and therefore unusable.

This is worth sitting with. The purchase is a workaround for the internal data problem, not a solution to a knowledge gap. And it's usually a worse signal than the real thing, because your own first-party behavioural data is specific to your customers and your products, while the vendor's is a population average.

The test before any purchase: could we derive this from data we already hold, if it were unified? If yes, you're deciding between buying a proxy forever and fixing the foundation once. Sometimes buying is genuinely the right short-term call — but make it a decision, not an accident.

Failure two: buying data you can't join

The second failure is more concrete. You buy an excellent property-attributes dataset keyed by precise address. Your policy records are geocoded to postcode centroid with inconsistent address formatting.

The join succeeds for the clean records — often 60–70% — and fails silently for the rest. Now your model has the enrichment for most risks and nulls for the remainder, and those nulls are not random: they correlate with older properties, rural addresses, and unusual buildings. Exactly the segments where you most needed the data.

Your enrichment has quietly introduced a systematic bias, and nothing in your monitoring will flag it.

The prerequisite for buying joinable data is the same one that keeps appearing: precise geocoding and resolved entity identity. Without it you're buying a dataset you can only partially use, at full price, with a hidden bias in the gaps.

A buying framework that avoids both

1. Name the decision first. Which decision changes because of this data? If nobody can answer crisply, you're buying reassurance. "It'll improve our models" is not a decision.

2. Check whether you already own it. Genuinely check. A surprising share of purchased signals are derivable internally once identity is resolved.

3. Test the join rate before you sign. Get a sample and attempt the join against your real records — not your cleanest subset. A 65% match rate changes the economics entirely, and vendors rarely volunteer it.

4. Analyse who falls out. Profile the unmatched records. If they cluster — older, rural, unusual — you've found a bias risk that must be managed, not ignored.

5. Measure incremental lift, not accuracy. The question isn't "is this data predictive?" It's "does it improve our model beyond what we already have?" Vendors demonstrate the former. Only you can measure the latter, and correlated signals often add almost nothing.

6. Confirm you may actually use it. Permitted use, consent, and fairness constraints. A signal that's predictive but acts as a proxy for a protected characteristic is a compliance problem you've purchased. Under the NAIC bulletin and EU AI Act, "the vendor said it was fine" is not a defence — you're the regulated entity.

7. Re-test annually. Vendor data quality drifts, your book changes, and yesterday's lift may be gone. Most insurers renew without ever re-measuring.

The uncomfortable summary

Third-party data is genuinely valuable, and some of it — property attributes, geospatial layers, weather history — fills gaps you'll never fill internally. Buy those.

But a meaningful share of third-party spend across the industry is one of two things: a paid workaround for unresolved internal identity, or a dataset that only partially joins and silently biases the models it feeds.

Both are symptoms of the same underlying gap. And both are the kind of spending that grows every year while the thing causing it goes unfunded, because fixing identity and geocoding has no vendor demo and no launch event.

Before the next renewal, run the two tests: could we derive this ourselves? and what's our real join rate? The answers usually reframe the budget conversation entirely.

We build the identity resolution, geocoding, and enrichment pipelines that make third-party data actually usable. More at IntelliBooks.

Comments

Popular posts from this blog

Why Your Insurance Data Warehouse Didn't Fix Anything

Straight-Through Processing: From 10% to 90%

Embedded Insurance: Why the API Is the Easy Part