Designs the tests that say whether a system is good enough to ship.
Why this one is hard to judge
Writing an eval is easy. Writing one that would actually have caught the failure is not.
What to ask for
Ask for an evaluation they designed and what it actually caught before release.
Ask how they chose the test cases, and who reviewed them for bias.
Ask for an example where a model passed their eval but still failed in use.
Ask how they decide when an eval result is good enough to ship.
The mistake most hirers make
Hirers treat a high score on a benchmark as proof of safety, without asking who wrote the benchmark. A candidate who can explain why their evaluation might be wrong is more trustworthy than one who is certain it is right. Numbers on a dashboard can hide the exact failures that matter most.
What good looks like after 90 days
An evaluation suite that reflects the specific ways the system could actually fail. At least one case where the eval caught a problem before users did. A clear, honest account of what the evaluation still cannot catch.
How we assess it
Against a rubric that is published in full, on evidence the practitioner supplies and a reviewer checks. Where something has not been verified, the profile says so.