Decides what to build, for whom, and how to tell whether it worked.
What separates them
The tell is what they decided not to automate; anyone can describe what they shipped, and it takes real judgement to explain what they deliberately left alone.
Ask these
01
Tell me about a feature you decided not to build, or a use case you decided not to automate, even though it was technically possible.
Probes judgement about limits, the thing that never shows up on a CV.
A strong answer
They give a specific example and state the actual reason, for example that the cost of an error was too high, or that human judgement genuinely outperformed the model for that case. They also say what they built instead.
The confident failure
They give a generic answer about keeping the user in mind without naming a specific case, or claim everything they scoped ended up shipping because it all made sense at the time.
02
How do you decide when a model is good enough to put in front of real users, versus when it needs more work?
Tests whether launch readiness is decided on evidence or on a feeling.
A strong answer
They name specific criteria tied to user harm and business impact, and describe a staged rollout or a threshold agreed with stakeholders in advance. They also reference a metric that had to clear before launch.
The confident failure
They say they test it thoroughly and use their judgement, with no specific threshold or process, so the decision sounds like a feeling rather than a gate.
03
Describe a time users behaved very differently to how you expected once a feature was live. What did you do?
Tests responsiveness to reality over attachment to the original plan.
A strong answer
They give a specific example of unexpected use and explain how they actually detected it, rather than just having heard about it anecdotally. They also describe a concrete change that followed.
The confident failure
They describe discovering unexpected use but nothing changed as a result, or the story is generic enough to apply to any product regardless of whether AI was involved.
04
How do you explain to a non technical stakeholder why an AI feature can't be made to work reliably 100% of the time, without losing their confidence in the whole project?
Tests communication of probabilistic reality to people who want certainty, an underrated product skill.
A strong answer
They describe framing the conversation around decision consequences and mitigations, such as a review step for high stakes cases, rather than technical caveats, and give a specific example of that conversation.
The confident failure
They give a technical explanation of model limitations that would go over a non technical stakeholder's head, or promise reliability improvements without a credible mechanism for delivering them.
What we ask when assessing for the register
Harder, and answerable only by somebody who has done the work. Published because a question that stops working when it is known was never testing anything.
01
Tell me about a product decision you made that you now think was wrong, where the model or the automation itself was not the actual problem. What was?
Tests self awareness about product failure beyond blaming the technology.
A strong answer
They give a specific example where the real issue was workflow fit or trust rather than model quality, and describe what they would do differently now.
The confident failure
They blame the model or the data, described as the technology wasn't ready, and avoid any admission of a genuine product misjudgement.
02
Describe how you scoped an AI feature differently from how you would have scoped an equivalent non AI feature. What actually changed in your process?
Tests whether product judgement is genuinely adapted to AI's particular failure modes or is standard process with AI vocabulary layered on.
A strong answer
They name specific adaptations, for example building in a fallback path for low confidence outputs, or agreeing acceptable error rates before the build started.
The confident failure
They describe a standard product process, discovery followed by prioritisation and a roadmap, with nothing specific to AI's uncertainty or failure characteristics.
Ask these whatever the discipline
Tell me about something you built that failed in production. What broke, how did you find out, and what did you change?
What would you refuse to do on this project, and what would you tell me instead?
How would you know, three months in, that this was not working?
What is the part of your own work that you are least confident about?
Relevant here, and none of them is evidence on its own. What each does and does not prove is set out in full on the certifications page.
Artificial Intelligence Governance Professional, AIGP, International Association of Privacy Professionals. That the holder has ever made a governance decision that cost somebody something. Governance is easy to describe and hard to enforce, and the exam can only test the first.
Foundation Certificate in Artificial Intelligence, BCS, The Chartered Institute for IT. Anything at all about building or assessing AI systems. BCS states there are no entry requirements, and the certificate is honest about being a foundation.