An engineering work sample that can’t leak.
Take-home tests get shared, and AI writes the code in minutes. Unfound tests what it can’t do for them: the judgment around a release when the deadline and the evidence disagree.
Debug a failing release under a deadline.
A release is due at five. A checkout test is failing intermittently, the on-call engineer has logs but no time, and the product manager wants to ship and fix forward. The candidate decides what to verify, what to ship and what to say.
Release is at 5pm. The checkout test is failing intermittently. Can we ship and fix forward?
What the rubric looks for.
Debugging approach
Narrows the problem with evidence, not guesses.
Risk judgment
Knows what a failure would cost before deciding to ship.
Communication under pressure
Tells the PM what they know, what they don’t, and when they will.
Prioritisation
Fixes the thing that matters first.
Every score, with its moment.
An illustrative excerpt of the evidence report for this role. Every number links to the decision behind it.
Held the release and asked on-call for the last three failing runs before deciding anything.
Told the PM the new ship time and the reason, instead of promising five o’clock.
Engineering hiring, explained.
Does it replace our coding interview?
No. It complements it. A coding exercise checks whether someone can write the code; the simulation checks the decisions around it: what to verify, when to ship, how to tell the team. Pair it with your own technical interview.
Which engineering roles does it fit?
Roles where release and incident decisions are part of the job: backend, platform and product engineers, and anyone who carries a pager.
Why can’t it leak like a take-home test?
Each run is generated for the candidate and forks at every decision they make, so there’s no fixed answer to share. What gets scored is the path they took, timestamped and replayable.

