Interview Coding Challenges Are Broken
Invert This Binary Tree, Please
I’ve sat on both sides of this table more times than I can count. Interviewing people, being interviewed, running loops, writing scorecards, arguing in debriefs about whether someone’s solution was “elegant enough”.
And I’ve slowly come round to an uncomfortable conclusion: the correlation between how someone does in a technical interview and how they do in the job is weak. Not zero. Weak.
We have built an industry-wide filter, applied it to nearly every hire we make, and mostly not checked whether it measures the thing we care about.
Two Lists That Barely Overlap
Here’s what a typical technical interview measures:
- Can you solve a contrived puzzle under time pressure, while a stranger watches?
- Have you memorised a specific set of algorithms and their edge cases?
- Can you perform, socially, in an artificial setting?
- Did you have several months of spare evenings to grind practice problems?
And here’s what the job asks of someone once they’ve started:
- Can you find your way around a large, messy codebase somebody else wrote?
- Can you explain a technical trade-off to a person who does not write code?
- Can you take a vague, half-specified request and turn it into a plan?
- Can you learn an unfamiliar domain quickly, including the boring business rules?
- Can you keep shipping reliable software, week after week, for years?
Read those two lists again. The overlap is smaller than our process assumes, and the second list is almost entirely absent from the first.

The Pattern I Keep Seeing
The failures I remember most clearly all interviewed beautifully.
They knew the algorithms. They talked fluently about complexity. They handled the follow-up question with the extra constraint. On paper, and in the debrief, they were the obvious hire. Then they arrived, and they couldn’t navigate a real codebase, couldn’t have a productive conversation with the people who actually needed the software, and couldn’t finish anything without someone sitting beside them.
Meanwhile some of the best people I’ve worked with had rough interviews. Nervous, rusty on tree rotations, or simply unlucky in which problem came up. Give them a real ticket in a real system and they were exceptional.
Both of these are individually explainable. Together, and repeatedly, they’re a hint that the test isn’t testing what we think.
The Grind Is a Filter, Just Not for Skill
Here’s the bit I find hardest to defend.
Doing well at leetcode-style interviews takes months of deliberate practice at a skill you will almost never use again. Which means what we are partly measuring is who had months of free evenings.
That is not evenly distributed. People with caring responsibilities don’t have it. People working long hours in a job they’re trying to leave don’t have it. People who’ve been out of the industry and are coming back don’t have it. We tell ourselves we’re selecting for raw ability and a fair amount of the time we’re selecting for available time.
The Argument Against
Now, the case for the current system, and it isn’t nothing.
Interviews have to be standardised. You cannot spend three days with every candidate, and you cannot make a hiring decision on vibes without discovering exactly how biased your vibes are. A consistent problem asked of everyone at least lets you compare like with like, and a puzzle with a known solution is easy to calibrate across a dozen interviewers.
Some signal also beats no signal. Weak correlation is still correlation. Someone who cannot write a loop under any conditions probably shouldn’t be writing production code, and a coding exercise does catch that.
And the alternatives all have their own problems. Take-home projects quietly demand unpaid evenings, which is the same fairness issue in different trousers. Pair programming sessions are better but hard to standardise and heavily dependent on the interviewer. Trial periods are the most accurate thing we have and almost nobody can afford to take one. It’s entirely possible that what we do now is the least bad of a bad set.
Where I Might Be Wrong
It’s also possible I’m simply not very good at this.
Interviewing is a skill, and a genuinely good interviewer probably extracts far more signal from the same forty-five minutes than I do. Every conclusion I’ve drawn above comes from watching my own hires, which is a small sample assessed by an interested party.
And “the correlation is weak” is a claim I’m making from experience rather than data. I believe it. I can’t prove it to you, and you should weight it accordingly.
A fair amount of the time we aren’t selecting for ability, we’re selecting for available time.
Both Sides of the Table
If you’re hiring:
- Use a realistic problem. A small bug in a real-ish codebase beats a puzzle every time
- Let people use their own editor, their tools, and the internet. That’s the job
- Weight communication as heavily as the code. You’re hiring a colleague, not a compiler
- Look for how someone learns, not what they’ve memorised
- Allow for nerves. Say so out loud, and give them a minute to settle
- Write down what the role actually needs before you design the test. Then test that
If you’re the one being interviewed:
- Practise anyway. The game is broken and you still have to play it
- Narrate your thinking. Half of what a good interviewer wants is your reasoning
- Ask what the process is up front. A company that has thought about it will tell you gladly
- Don’t read too much into a rejection. The signal is noisy in both directions
The system is a bit rubbish and we’re all stuck in it together. That’s not a reason to run yours the way it was run on you.

Until next time, happy coding!