Skip to content
Over Unity

Interview questions

MLOps and platform

Owns the pipelines, deployment, monitoring and cost of models in production.

What separates them

This work is judged by the absence of incidents, which is invisible right up until it stops being true.

Ask these

01

Tell me about the worst outage you have dealt with on a platform you ran. What was the actual root cause?

Reveals real operational depth as opposed to theoretical infrastructure knowledge.

A strong answer

Gives a specific technical root cause, a timeline, and a structural change made afterward, not just a description of fixing a bug.

The confident failure

Describes the incident in vague terms and the fix is simply adding better monitoring, with no further specifics.

02

How do you know your training and serving infrastructure is healthy before someone tells you it is not?

Tests whether they build for observability rather than reacting once users complain.

A strong answer

Describes specific leading-indicator metrics and alerts, and gives an example of catching an issue before it reached users.

The confident failure

Says they have dashboards for everything, without naming what actually triggers action on any of them.

03

Tell me about a piece of infrastructure you built that you later had to rebuild because it did not hold up.

Tests honesty and whether they learn from hitting real scaling limits.

A strong answer

Names a specific technical limitation reached, what triggered the decision to rebuild, and what changed in the second version.

The confident failure

Describes only successful builds and cannot name anything that needed reworking.

04

How do you handle a model or pipeline that needs to roll back quickly after a bad deployment?

Tests whether rollback is designed in advance or improvised under pressure.

A strong answer

Describes a concrete rollback mechanism and versioning approach, and gives a real time-to-recover figure from an actual incident.

The confident failure

Says they can always roll back, without describing how, or citing any real instance where it was needed.

What we ask when assessing for the register

Harder, and answerable only by somebody who has done the work. Published because a question that stops working when it is known was never testing anything.

01

Describe the cost of an infrastructure decision you made that turned out to be wrong, and how you found out.

Tests depth of real experience, since infrastructure trade-offs often only reveal their true cost later.

A strong answer

Names a specific decision, such as a storage choice or a scaling strategy, what it cost in money or time, and how it was corrected.

The confident failure

Describes a hypothetical trade-off analysis in the abstract without ever naming a decision that actually went wrong.

02

Walk me through how you would design monitoring for a system where silent failure is more dangerous than a crash.

Tests understanding of the hardest platform problem, since a crash is easy to see and quiet degradation is not.

A strong answer

Describes specific techniques such as running a new version alongside the old one and comparing outputs statistically, with a concrete example.

The confident failure

Says they would add logging and alerts, without addressing how to detect a failure that still produces plausible output.

Ask these whatever the discipline

  • Tell me about something you built that failed in production. What broke, how did you find out, and what did you change?
  • What would you refuse to do on this project, and what would you tell me instead?
  • How would you know, three months in, that this was not working?
  • What is the part of your own work that you are least confident about?
With the answer patterns

If they hold a certification

Relevant here, and none of them is evidence on its own. What each does and does not prove is set out in full on the certifications page.

  • AWS Certified Machine Learning Engineer, Associate, Amazon Web Services. That the holder has run anything at scale, or that they can tell a good architecture from a defensible one. It is scoped to one cloud, and the exam rewards recognising the AWS answer rather than the right answer.
  • Professional Machine Learning Engineer, Google Cloud. Anything about judgement on problems Google Cloud has no service for, which is most of the interesting ones.
  • Azure Data Scientist Associate, DP-100, Microsoft. Statistical judgement. Nothing in the syllabus tests whether the holder would notice a leak between training and evaluation data.
  • Certified Machine Learning Professional, Databricks. Transferability. It is the most platform-bound certification on this list.
  • Deep Learning Institute certifications, NVIDIA. Independent judgement about whether the workload needed that stack in the first place.
The full register
Or let us assess themWhat this work pays

Over Unity makes introductions between hirers and independent specialists. It is not a party to any engagement, does not hold or transfer payments, and does not determine employment status. Specialists are never charged a fee.