Designs and ships products where a language model is the core of the product.
Why this one is hard to judge
The gap between a working prototype and something that survives real users is where most of the work is, and it does not show in a portfolio.
What to ask for
Ask for an app they shipped and how they know whether users trust its answers.
Ask what their evaluation set looks like and who wrote the test cases.
Ask for an example where the model gave a wrong answer confidently, and what they changed.
Ask how they decided between a bigger model and a better prompt.
The mistake most hirers make
Hirers get impressed by a slick demo and skip asking how the person tests for wrong answers. A demo that works once for you is not evidence the app behaves well for a stranger with a messy question. The gap between demo and dependable product is where this discipline actually lives.
What good looks like after 90 days
A feature used by real users with a written evaluation set behind it. A clear record of failure modes found and fixed, not just launched. A sensible view on when to change the prompt versus the underlying model.
How we assess it
Against a rubric that is published in full, on evidence the practitioner supplies and a reviewer checks. Where something has not been verified, the profile says so.