70,260 configurations: what it takes to build reliable insurance AI

What happens when nine insurer plan files turn into 70,260 valid configurations and more than 8.2 million comparison cells?

Allan Cândido

Marketing Executive, Cluda

Nine insurer plan files can look like a relatively straightforward comparison exercise. Line up the benefits, identify the differences, work through the options relevant to the client.

In practice, those nine files represented 70,260 valid product configurations and more than 8.2 million individual comparison cells.

That scale isn't the interesting part. For AI policy comparison, what matters is whether the system remains reliable across all of them — and whether every answer can be verified rather than inferred.

Nine files are not nine comparisons: why reliable AI policy comparison is harder than it looks

A benefits or commercial policy comparison isn't a side-by-side read. In practice, each plan can be configured in multiple ways before a comparison even begins — underwriting basis, excess, benefit level, hospital list and optional cover can each introduce another set of choices, and they vary insurer by insurer and product by product.

Take a relatively simple example: six underwriting options, nine plan levels and one binary optional benefit already produce 108 valid configurations. Apply that logic across the nine plan files in our test set and the total becomes 70,260 distinct configurations. Render a comparison for every one of them and you get 8,269,008 individual cells to validate.

In a live renewal, a broker only needs to work through the configurations relevant to that client. The remaining possibilities are rarely checked individually, simply because the underlying search space is too large to review manually.

Once you have that many valid configurations, you need every underlying data point checked: the limit, the excess, the waiting period, the exclusion wording, the definition of "pre-existing condition" as that specific insurer defines it. The scale itself is the headline, but it's the wrong one. Any system can produce a large number by multiplying options together. The question that actually matters for a broker relying on the output is narrower: when it says a cell contains a particular limit or exclusion, can that claim be traced to a specific point in a specific insurer's published document — or is it an inference, a best guess, or an average of what that insurer "usually" does?

The number that matters isn't the number

Every one of those 8,269,008 cells in our test was checked against the expected result for that specific configuration — not an inference based on typical terms for that insurer, not something consistent with prior wordings we'd seen. The resulting values remain traceable to the relevant source documentation.

That distinction sounds pedantic until the alternative shows up in a client conversation. This is a structurally difficult problem: any process, manual or automated, that fills a gap with a plausible assumption instead of a checked answer can still produce plausible results most of the time. The problem is that the exceptions may be difficult to spot until they matter. They surface at claim, when the schedule says one thing, the wording says another, and nobody can reconstruct why the client was told what they were told.

Traceability is what turns "we compared these policies" into something a broker can actually stand behind. It's also, increasingly, what the regulator is asking to see.

This also matters for Consumer Duty

The FCA's focus has increasingly moved towards whether firms can evidence the analysis behind customer outcomes, rather than simply state that appropriate processes exist — its own guidance is explicit that fair value assessments "should reflect the analysis undertaken as part of these processes, rather than being produced separately and/or retrospectively" (FCA, Price and Value Outcome: good and poor practice), and firms are now expected to "demonstrate — not merely claim" good customer outcomes (Browne Jacobson, FCA anticipated priorities for insurance brokers and intermediaries, 2026). Three years into Consumer Duty, showing the working matters as much as reaching the right answer (Clifford Chance, The Consumer Duty at Three, May 2026). For comparison work specifically, that makes a contemporaneous audit trail considerably more useful than one reconstructed after a decision has already been made.

A spreadsheet comparison can still produce a defensible recommendation. What's harder to produce, without significant manual effort, is a record showing which clause supported which conclusion for every configuration a client could have been offered.

Four questions worth asking about any comparison you rely on

Whether it's a manual process, a spreadsheet template, or a piece of software, the same checks apply before you trust the output on a live client case:

Can every stated data point be traced back to its source? If the answer is "it's usually right" rather than "here's the source," that's a gap worth knowing about before a claim exposes it.

Does it tell you when it doesn't know, or does it guess? A tool or process that fills silence with a plausible assumption is harder to rely on than one that visibly flags a missing data point, because the first kind fails silently.

Would the audit trail exist before someone asked for it? Contemporaneous beats retrospective, per the FCA's own guidance above. If the evidence only gets assembled when a file is pulled for review, it isn't really an audit trail.

How many configurations were actually checked versus assumed consistent? Two or three spot-checked scenarios standing in for tens of thousands of valid configurations is a reasonable practical compromise — as long as everyone involved knows that's what happened, rather than believing the comparison was exhaustive when it wasn't.

What we tested, and what it's worth to a broker

The 70,260-configuration exercise was an internal stress test of Cluda's own comparison engine, run against nine live insurer plan files rather than sample data, specifically to check that traceability held at scale and not just on the handful of scenarios a demo would show. The complete run took 172 seconds of machine time. We ran every valid configuration and checked every rendered cell against the expected result, rather than accepting anything because it looked "probably right" or consistent with a clause seen elsewhere. That is the standard we set for our own comparison engine: not because the number of combinations is impressive, but because traceability has to hold when the system moves beyond the handful of scenarios that would normally appear in a demo.

The broker should never need to think about 70,260 configurations individually. That is precisely the point. The complexity belongs in the system, so the broker can concentrate on the comparison in front of them.

The goal, at every stage, is reliable delegation: work that has genuinely moved from person to system, checked in a way that holds up under scrutiny, with responsibility never in doubt about where it sits. Every new level of delegation requires a new level of control. Firms that treat that as the rule, not the caveat, are the ones that will still be able to explain their AI use with confidence in two years' time.

Frequently Asked Questions

Frequently Asked Questions

What makes a policy comparison "traceable"?

Every stated data point links back to its source in the actual insurer document reviewed, not to an assumption about what that insurer "usually" includes. If a tool or process can't point to the source for an answer, it becomes harder for the broker to distinguish a verified fact from an assumption.

Does the FCA require brokers to document how a comparison was reached?

Not in those exact words, but its Consumer Duty guidance on fair value is explicit that assessments should reflect the analysis as it happened, not be reconstructed afterwards. In practice, that means the record needs to exist at the time the comparison was made.

What's the risk of trusting a comparison that can't show its sources?

The gaps don't announce themselves. A plausible assumption standing in for a checked fact looks identical to a verified one, right up until it surfaces at claim and nobody can reconstruct why the client was told what they were told.

How many configurations should actually be checked before recommending a policy?

There's no fixed number — it depends on the client and the risk. What matters is knowing which configurations were checked against source wording and which were assumed, rather than treating every comparison as exhaustive when it wasn't.