Builds agentic workflows and retrieval systems that return the right thing.
Why this one is hard to judge
Retrieval quality is measurable, but almost nobody measures it. Ask for the numbers and most cannot produce any.
What to ask for
Ask for a specific case where an agent looped or took the wrong action, and how they caught it.
Ask how they measure whether retrieved documents were actually relevant.
Ask what guardrails stop an agent from taking an irreversible action by mistake.
Ask them to describe a task they decided not to automate with an agent.
The mistake most hirers make
Hirers assume an agent that works in a demo will behave the same way with real, messy data and real stakes. They do not ask what stops the agent from doing something irreversible. The candidate who lists the failure modes honestly is usually the safer hire than the one with a slicker demo.
What good looks like after 90 days
One agent or retrieval system running with clear limits on what it can do unsupervised. A written list of failure modes observed and how they are caught. At least one task they recommended not automating, and why.
How we assess it
Against a rubric that is published in full, on evidence the practitioner supplies and a reviewer checks. Where something has not been verified, the profile says so.