Designs and ships products where a language model is the core of the product.
What separates them
The real work starts after the prototype works, when real users do things the demo never anticipated.
Ask these
01
Tell me about a feature powered by a language model that you had to significantly change after real users touched it.
Tests whether they have shipped and iterated on a real product rather than only built demos.
A strong answer
Describes a specific user behaviour that broke an assumption, the concrete change made, and how the improvement was actually measured.
The confident failure
Tells a polished launch story where users loved it, with no specific friction or change described anywhere.
02
How do you handle the case where the model gives a plausible but wrong answer inside your product?
Tests their approach to hallucination and reliability, which is the core problem in this discipline.
A strong answer
Describes specific guardrails, verification steps or interface patterns for showing uncertainty, with a real example of one implemented.
The confident failure
Says a good system prompt prevents that, offered as the entire answer with nothing else in the safety chain.
03
What have you had to remove or simplify from an LLM product because it did not hold up with real users?
Tests honesty about failure and the understanding that more capability is not always better.
A strong answer
Names a specific feature that was cut, the user signal that drove the decision, and what replaced it.
The confident failure
Gives a vague answer about iterating based on feedback, without naming any specific feature that was actually removed.
04
How do you decide what to show the user when the model is uncertain about its own answer?
Tests design judgement specific to probabilistic systems, which most product experience does not cover.
A strong answer
Describes concrete confidence signalling or fallback flows, drawn from a real product decision they made.
The confident failure
Says they just let the model answer, or describes a generic disclaimer with no real design thought behind it.
What we ask when assessing for the register
Harder, and answerable only by somebody who has done the work. Published because a question that stops working when it is known was never testing anything.
01
Tell me about a metric you used to decide an LLM feature was actually working, beyond user sentiment.
Distinguishes people who measure real outcomes from those relying on anecdote and impression.
A strong answer
Names a concrete metric tied to a business outcome, such as task completion or correction rate, and explains exactly how it was measured.
The confident failure
Cites satisfaction scores or engagement, without ever connecting that number to whether the task was actually completed correctly.
02
Describe a case where a prompt change improved one part of the product and quietly broke another.
Tests understanding of regression risk in prompt-based systems, which is often invisible without deliberate checking.
A strong answer
Describes a real example, how it was caught, such as an evaluation set or monitoring, and what changed in the process afterward.
The confident failure
Says they test carefully before every change, without naming a method or being able to point to a real instance where it failed.
Ask these whatever the discipline
Tell me about something you built that failed in production. What broke, how did you find out, and what did you change?
What would you refuse to do on this project, and what would you tell me instead?
How would you know, three months in, that this was not working?
What is the part of your own work that you are least confident about?
Relevant here, and none of them is evidence on its own. What each does and does not prove is set out in full on the certifications page.
Azure AI Engineer Associate, AI-102, Microsoft. That the holder can evaluate whether a retrieval system is returning the right thing. Assembly and evaluation are different skills, and the second one is the one that is scarce.