Skip to content
Over Unity

Insights/For specialists

How to explain model risk to a non-technical client

7 minute read. Updated 2026-08-08.

The short answer

The analogies that survive scrutiny with a non-technical client describe a model as a confident junior who never says 'I don't know', and production monitoring as spot checks on a line with an agreed sampling rate. Retire 'black box' and any comparison to human accuracy, both mislead. Sequence the conversation as strengths, failures, downstream consequences, then mitigation, and get the acceptable failure rate agreed in writing before go live.

Why this conversation decides the contract

A CTO who signs off a model they don't fully understand isn't signing off on your code. They're signing off on the 20 minutes in which you explained what could go wrong and what you'd do about it. That's the actual deliverable in the first meeting, whatever the statement of work says the deliverable is.

Get this wrong and the work that does land gets micromanaged, because the client never trusted you to see the failure coming before it happened. Get it right and you get called back for the next model without a fresh procurement round. Treat it as a skill worth rehearsing as deliberately as your architecture choices, because it decides more contracts than either does.

The confident junior

The analogy that survives scrutiny: the model is a new starter who has read almost everything ever published on the subject but was never taught to say 'I don't know'. It answers a question it has never encountered in exactly the same confident tone as one it has answered 1,000 times before. Fluency doesn't tell you which is which, and neither does the client's gut.

This holds up because it names the failure mode instead of softening it, and it leads the client straight to a fix they already understand. Nobody lets a new starter approve a six-figure invoice alone in their first month either. You put a review step around the decisions that matter and widen it as trust is earned, and that's a governance conversation most clients have already had with human staff.

Push it one step further and you answer the question before they ask it: what happens when the junior is confidently wrong and nobody checks? That's where you describe the escalation path in plain terms, not as an architecture diagram, but as who sees the answer before it reaches the customer.

Spot checks on a production line

The second analogy worth keeping is quality control on a production line, not a single machine. Nobody inspects every unit that comes off it. You sample at an agreed rate, you know what happens when a sample fails, and you keep a record of both. That is what monitoring a model in production actually is.

It works commercially because it doesn't promise zero defects, which no honest supplier promises on a physical line either. It gives the client something to negotiate: how often you check, what it costs when a fault slips through, who's accountable for missing it. A COO has had that exact conversation before, about claims processing or manufacturing, so nothing about it feels new to them.

The analogy also protects you. Once the sampling rate and the escalation path are agreed in writing, a failure that falls inside the agreed rate is not a broken promise. It's the system working as specified. That distinction matters enormously the first time something goes wrong in production.

Two analogies to retire

'Black box' is the one every candidate reaches for, and it's the one to drop first. It implies the system is unknowable, and an unknowable system sounds like nobody is accountable for what comes out of it, which is the exact impression you don't want to leave in a risk conversation. Say instead that you can't see the model's internal reasoning, but you can test its behaviour at the boundary, the way an electrician tests a circuit without opening the wall: known input, observed output, repeated until the pattern is reliable.

The second to retire is any comparison to human accuracy along the lines of 'it's right more often than a person would be'. Even without a figure attached, the comparison misleads, because it hides how the two fail differently. A team of people makes independent mistakes that mostly cancel out across the group. A model making the same call 1,000 times a day fails the same way 1,000 times, and those failures are correlated, not scattered.

That difference is the entire reason review processes built for AI systems look different from review processes built for people. Glossing over it with a flattering comparison buys you a good first meeting and costs you credibility the day something goes wrong, because you'll have already told the client the risk was smaller than a person's.

The order the conversation should run in

Sequence matters as much as content. Open with what the system does well, and be specific about it, or the client hears nothing that follows a disclaimer. Move to where it's known to fail, using whichever analogy fits the failure. Then describe what happens downstream when it does fail, inside the client's process, not yours. Close with what you're doing to catch it before it reaches a customer or a regulator.

Never promise zero risk, and say so out loud rather than letting silence imply it. What you can promise is an agreed failure mode: the client signs off in writing, before go live, on what an acceptable rate of error looks like for this system. That single document protects you more than any technical excellence, because it moves the argument from 'was this good enough' to 'did we meet what we agreed'.

What this earns you

None of this is a soft skill bolted on after the technical work is done. It's the reason a client renews without running a fresh procurement round, and the reason they name you when a colleague at another company asks who they used.

Rehearse the two analogies until they come out as plain sentences, not lecture notes, and retire the two that will get you caught out under questioning. The client won't remember your model's evaluation score. They'll remember whether you told them the truth about what could go wrong, and that's what gets you the next contract.

What to do about it

  • Use the confident junior analogy to explain hallucination without minimising it.
  • Use the production line analogy to explain monitoring, and agree a sampling rate in writing.
  • Retire 'black box'; describe testing the model's behaviour at the boundary instead.
  • Never compare model accuracy to human accuracy, even without a number attached.
  • Sequence the conversation: strengths, failure modes, downstream consequences, mitigation.
  • Get the client to sign off an acceptable failure rate in writing before go live.

Questions people also ask

Isn't 'black box' just an accurate description of deep learning?

It's technically defensible but commercially costly. It implies nobody, including you, can be held accountable for the output, which is the opposite of what a risk conversation needs to establish. Describe instead what you can do: test the model's behaviour at defined boundaries with known inputs, even without seeing its internal reasoning. That gives the client something to act on, which 'black box' never does.

How do I explain hallucination without frightening the client off the project entirely?

Use the confident junior analogy. It names the failure honestly, a fluent, wrong answer delivered with no hedge, without implying the whole system is unreliable. It also leads naturally to the fix the client already understands from managing people: review the parts that matter, and widen trust as the record improves. Naming the risk and describing the control together is what keeps the conversation moving rather than stalling it.

What do I say if the client asks how often it will fail?

Don't invent a figure to sound precise. Agree, in writing, what an acceptable rate of error looks like before the system goes live, based on what you observe in testing and what the client's process can tolerate downstream. That document is what protects both of you later, because it turns 'is this good enough' into 'did we meet what we agreed', which is a much shorter argument to have.

Should I use technical terms like precision and recall in this conversation?

Only after you've made the point in plain language first. A CTO or COO can learn what precision and recall mean, but the first meeting isn't the place to teach them while also asking them to sign off risk. Lead with the analogy, get agreement on the principle, then introduce the technical term as a label for something they've already understood, not as the explanation itself.

Where the figures come from

Every rate and salary quoted in this article is a median or percentile of figures advertised in UK job postings over the six months to 8 August 2026. They are not rates paid, and the gap widens at the top of a range.

The full salary guide, with sample sizes

More for specialists

Apply to the registerWhat we are assessing for

Over Unity makes introductions between hirers and independent specialists. It is not a party to any engagement, does not hold or transfer payments, and does not determine employment status. Specialists are never charged a fee.