The engineering test for the AI era

Every candidate is using AI. The best ones know when it's wrong.

An AI-native technical interview that measures engineering judgment, not AI output.

The test

A technical interview for the modern day, with a twist.

01 · Build

A real engineering task, not a coding puzzle.

Candidates receive a realistic product brief and a fixed amount of time to build. They work in an environment that feels like their day-to-day workflow, with an AI coding agent available from the start.

02 · Collaborate

An AI agent that's helpful. Well, most of the time.

Just like Codex, Cursor, or Claude Code, the agent can propose designs, write implementations, and answers questions. But some suggestions are intentionally flawed, revealing whether candidates verify, challenge, or blindly accept AI output.

03 · Measure

Engineering judgment becomes the signal.

We evaluate more than the final project. We track how candidates used the AI, when they questioned its recommendations, what they corrected, and how effectively they balanced speed with sound engineering decisions.

04 · Hire

Clear evidence instead of gut feeling.

Interviewers receive a concise report covering project progress, AI usage, critical decisions, and demonstrated engineering judgment—making it easier to identify candidates who can work effectively with AI instead of simply relying on it.

The mismatch

Your screen tests one thing. The job now demands another.

What the old screen measures
  • ×Recalls an algorithm cold
  • ×Writes syntax from memory
  • ×Solves a clean, closed puzzle
  • ×Works alone, no tools
What the job now demands
  • Directs an AI agent
  • Catches confident-wrong code
  • Turns a vague ask into the right build
  • Stands behind every shipped line
What you're actually hiring for

Three things the new job demands.

01 / Direction

Can they steer it?

AI coding assistants don't understand the nuances of your business. Modern engineers need to direct agents towards the best solution to specific problems, not just generic solutions.

How we prove itWe hand them a live agent which intentionally steers in sub-optimal directions. Do they push back, or ship the first thing it proposes?

02 / Judgment

Can they catch it?

AI writes code, and suggests ideas that look right but often aren't. Catching what's confidently wrong before it ships is what separates an engineer from a code generator.

How we prove itOur agent plants subtle, realistic bugs on purpose, and suggest subpar ideas intentionally. Your report shows exactly which they caught and which slipped through. You get a live replay of how it all happened.

03 / Ownership

Can they stand behind it?

Engineers must be accountable for every line that ships, whether they wrote it or the AI did. They must understand the why as much as the what.

How we prove itThey're graded on the final code as their own. Then, at the end, a neutral examiner asks them to defend the decisions they made.

The defense

The why, not just the how.

After the build, an AI examiner takes over. It reviews the finished project and generates questions tailored to the candidate's implementation, and your preferences, giving them a chance to explain their decisions and demonstrate their reasoning.

  • Grounded in their work. Every question comes from the code they actually shipped and the mistakes they did or didn't catch.
  • Tuned to your preferences. Care to ask mostly about the higher level architecture? We'll prioritize that. Testing strategy? Ins and outs of the code? Collaboration with the agent? What you're looking to learn shapes the interviewer's prompt.
  • Optional and tailored. Turn it on or off per role, change the time allowed per question, or conditionally trigger the entire section based on their part one performance.

Illustrative exchange — not a real candidate.

The output

See how they think, not just what they typed.

Measure both how well they managed the agent, as well as their final output. Supervision is measured against mistakes we planted on purpose. Craft is a read on how they build. See how many of the agent's mistakes they caught, and how well they met the spec they were charged with crafting all in one report.

Illustrative example — not a real candidate.

The stakes

Every hire is a six-figure swing.

Bad hires are more than just a mistake
One bad hire · fully loadedUSD
Mis-hire cost~30% of first-year earnings — US Dept. of Labor$36,000
Recruiting & backfillaverage cost to source and re-hire$6,200
Cost of the seat sitting empty$500/day until the next hire lands$22,500
Total per bad hire$64,700

Sources: US Dept. of Labor (≈30% of first-year earnings) · recruiting/backfill average · $500/day vacancy cost. Adjust the sliders to your own numbers.

drives the 30% mis-hire line
× $500/day of lost output & coverage
The flip side — what a great hire is worth
$100K–$200K+
in extra value over a typical two-year tenure

The spread between a top and a bottom hire runs ~$106,000 a year in output — every year they stay.

Book a demo

Get access to the platform and try it for yourself.

Please enter your name.
Please enter a valid work email.
Please tell us where you work.
Prefer email? [email protected]
Thanks — you're on the list. We'll reach out shortly to schedule your demo and share a sample report.