How to run a technical assessment for AI work without wasting anyone's time
7 minute read. Updated 2026-08-08.
Cap technical take-home tests at 2 hours. Longer tests filter for whoever has free time that week, not whoever is best, because your strongest candidates are already billing elsewhere at market rates. Replace the marathon exercise with a paid, capped task built on a real bug, a live walkthrough of the candidate's own past work, and a reference from someone who has actually reviewed their code.
The 8-hour take-home tests patience, not competence
Many AI hiring processes still include a take-home exercise that takes half a day or more to complete properly: a small model pipeline, a written report, code with tests. The intention is reasonable. The result rarely is.
A test that takes 8 hours filters for people who have 8 free hours. Your best candidates rarely do. They are contracted, they are billing, and their evenings belong to their own clients or their own families. The candidate who submits a polished 8-hour take-home may simply be the one currently between engagements.
This is not a claim that busy people are always better. It is a claim that a long take-home selects for availability before it selects for skill, and availability is not what you are hiring for.
What the exercise actually costs, on both sides
An AI specialist working at advertised UK contract rates near the reported median, £551 a day for artificial intelligence roles broadly and £575 a day for machine learning engineers over the six months to August 2026, is not idle for the rest of their week. Every unpaid hour on your test is an hour not billed elsewhere.
On your side the cost is less visible but larger. Reviewing take-home submissions properly, catching the candidate who has bolted together tutorial code rather than reasoned through the problem, takes real time from a senior engineer. Multiply that by a shortlist of 6 candidates and you have spent a full day of your best person's week screening people you will mostly reject.
The exercise looks free because no invoice arrives. It is not free.
Cap it at 2 hours, and mean it
2 hours is enough to see how someone reasons about a real problem. It is not enough to reward whoever grinds hardest through boilerplate. Set the cap, tell the candidate the cap, and design the task to fit inside it comfortably rather than to just about fit if nothing goes wrong.
Give them a genuine fragment of a real problem rather than a toy dataset stripped of context: a short section of a pipeline with a bug you have already found, a model output that is wrong in a specific and diagnosable way, a set of prompts producing an inconsistent result. Ask them to find it and explain it, not to build something from a blank page.
If a task cannot be attempted properly inside 2 hours, it is not a technical assessment. It is unpaid work dressed as one.
Run this instead of the marathon
Replace the long take-home with a live session built around the candidate's own past work. Ask them to walk you through a project they actually shipped: what broke, what they changed, what they would do differently now. This shows judgement, and judgement only shows up once something has gone wrong and someone has had to fix it.
Follow it with a short live technical conversation on your own ground: describe a real decision your team is facing, including the constraints that make it hard, and ask them to think out loud. Watch how they handle not knowing something. Nobody in this field knows everything, and the ones who pretend otherwise are the ones to worry about.
Close with a technical reference, not a character reference from a line manager but a conversation with someone who has actually reviewed this person's code or model output. Ask that person a direct question: would you hand this candidate a production system unsupervised. The answer carries more information than any scored test.
Pay for the 2 hours if you can
If you set a capped exercise, offer to pay for it, even a modest amount. Set against the day rates you will eventually be paying, the cost is small, and it signals that you understand the value of the time you are asking for.
It also changes who applies. An unpaid, lengthy take-home selects for people with time on their hands. A short, paid, respected exercise selects for people who are good enough to be busy and choose to spend 2 hours of that time on you.
Where a longer test is defensible
There is one case where more time is genuinely warranted: a role where the day-to-day work is itself slow and reflective, such as a red-teaming engagement where the whole point is patient, adversarial probing rather than fast diagnosis. Even then, structure it as paid work on a real system, not as an unpaid exam.
Outside that case, if you find yourself defending an 8-hour exercise on the grounds that the role demands rigour, ask honestly whether the rigour is in the task or in the length. Usually it is the length doing the work, and length is the easiest thing in the world to fake with enough free time.
Worth remembering too that narrow specialist categories such as generative AI, large language model work, prompt engineering and red teaming are new enough that the advertised market only reports a single median day rate near £550 for several of them, with no published spread. The talent pool is thin. Asking it to clear an 8-hour hurdle before you will even talk to them narrows it further, for no gain.
What this buys you
A capped, live, paid process takes less calendar time than the take-home model most companies still run, and it produces a better decision. You see how someone thinks under mild pressure, how they talk about a genuine failure, and how another practitioner rates their actual work.
None of that is knowable from a polished document submitted at midnight by someone who happened to have a free evening.
What to do about it
- Cap every take-home exercise at 2 hours and say so in writing.
- Base the exercise on a real, specific bug or inconsistency, not a green-field build.
- Pay for the time if you can, even a token amount.
- Replace long take-homes with a live walkthrough of the candidate's own past project.
- Get a technical reference from someone who has actually reviewed their code, not a manager reference.
- Treat a demand for rigour as a reason to check whether the task or the length is doing the work.
Questions people also ask
Isn't a take-home fairer than an interview because it removes bias?
A capped live session with a fixed rubric removes as much bias as a long take-home, without the time cost. A live conversation also reveals reasoning that a polished submission can hide, including work that was quietly outsourced or copied. Fairness comes from a consistent structure applied to everyone, not from the length of the exercise, so keep the structure and cut the hours.
What if the candidate refuses to do any test at all?
A senior specialist billing at advertised market rates has every reason to push back on unpaid, lengthy tests, so refusing one is not automatically a red flag. Ask instead for a walkthrough of past work or a short paid session. Most credible practitioners will happily do that. If they will not engage with any form of assessment, that is worth noting.
How do I compare candidates fairly if the live session goes differently each time?
Use the same core problem for everyone, even if the conversation branches differently from there. Score against a small number of things you are actually checking, such as whether they ask the right diagnostic questions, rather than whether they land a specific correct answer, because for most real AI problems there isn't one fixed answer.
Should the hiring manager or a technical peer run the live session?
Ideally both, in separate conversations. A hiring manager can assess communication and fit. Only a technical peer can tell whether the reasoning actually holds up under follow-up questions, and that is the piece a long take-home was originally meant to test.
Where the figures come from
Every rate and salary quoted in this article is a median or percentile of figures advertised in UK job postings over the six months to 8 August 2026. They are not rates paid, and the gap widens at the top of a range.
- IT Jobs Watch, UK contract rates, 6 months to 8 August 2026, read 2026-08-08.
- IT Jobs Watch, UK permanent salaries, 6 months to 8 August 2026, read 2026-08-08.