The Underwriter Workbench: Designing AI Experts Actually Use
Here's a pattern that repeats across the industry and rarely gets written about honestly. An insurer builds an AI underwriting tool. It's accurate. It's been validated. It goes live. And six months later, underwriters are still doing it the old way, using the tool only when someone's watching.
The project is recorded as delivered. The value never materialises. And the post-mortem, if there is one, blames "change management" — which is a way of saying "the users were the problem" without saying it out loud.
They usually weren't.
Why experts reject accurate tools
1. It doesn't show its reasoning. An underwriter is professionally accountable for the decision. A score with no explanation asks them to stake their judgment on something they can't interrogate. Refusing that isn't stubbornness — it's exactly the professional caution you hired them for.
2. It's confidently wrong on the cases they know best. Every experienced underwriter has a segment they understand deeply. If the tool is obviously wrong there — and models trained on thin data for unusual risks often are — they'll rationally distrust it everywhere else. One visible failure in their area of expertise costs you the whole tool.
3. It adds a step instead of removing one. If using the tool means opening another system, re-keying data, and interpreting an output, you've added work to a role already under time pressure. Accuracy doesn't compensate for friction.
4. It's aimed at the wrong part of the job. Tools frequently target the decision — the part the underwriter finds valuable and enjoys — while leaving untouched the genuinely tedious part: assembling and checking the information. That's precisely backwards.
5. Nobody asked them. The tool was specified by people who don't do the work, validated against historical data rather than against how decisions are actually made, and delivered as a finished thing.
What "workbench" means, and why the framing matters
The shift is from building a decision-maker to building a workbench — an environment that makes the expert dramatically faster at their existing job, with AI doing preparation rather than judgment.
Concretely, a submission arrives and the workbench has already:
- Extracted the data from the submission documents and schedules
- Enriched it — geocoding, property attributes, company data, prior claims
- Flagged the missing fields and the internal inconsistencies
- Surfaced comparable risks and how they were priced
- Highlighted the specific characteristics that make this risk unusual
- Produced an indicative score with its drivers shown
The underwriter then does what they're actually for: judge. In minutes rather than hours, on prepared information instead of raw documents.
Notice what this does to adoption. It removes work they dislike and leaves the work they value. Nobody has to be persuaded to use it — and the AI still runs on every submission, just aimed at preparation.
Design rules that determine whether it's used
Show the drivers, always. Never a bare score. "Elevated because: roof age 22 yrs, two prior claims, flood zone AE." An underwriter who can see the reasoning can agree, disagree, or spot that a driver is based on bad data — which is itself valuable.
Make disagreement first-class. Let them override, and capture why in a structured field. This gives you dignity for the expert, an audit trail for the regulator, and the single best training signal you'll get for improving the model. Systems that make overriding awkward lose both the trust and the data.
Be honest about confidence. The tool should say "I have little comparable data for this risk" rather than producing a confident number. Admitting uncertainty on the unusual cases is what earns trust on the routine ones.
Live where they already work. Inside the existing workflow, not another tab and another login.
Build it with them, iteratively. Not requirements gathering — actual co-design with the people who'll use it, shipping small and adjusting. They know the edge cases your historical data doesn't contain.
The measure that matters
Track adoption and time-to-decision, not just model accuracy. A model at 85% accuracy used on every submission creates enormous value. A model at 93% that nobody trusts creates none.
And watch the override rate as a health signal: consistently high in a segment means either the model is weak there or a driver is based on bad data. Either way it's telling you something specific and actionable — which a model deployed without a feedback path would never surface.
The point
Most failed insurance AI projects didn't fail technically. They produced a working model that experts declined to rely on, for reasons that were entirely rational from where those experts sat.
Aim the AI at the preparation, show your reasoning, make disagreement easy and useful, and build it with the people who'll use it. Do that and adoption stops being a change-management problem — because you've built something people actually want.
We build the data foundations and decision-support systems that experts actually adopt. More at IntelliBooks.
Comments
Post a Comment