From the People track

Interview loops that predict performance

Most loops measure interview skill. Give each stage exactly one job, and let the highest-weight stage be a structured judgment conversation with the people who'll actually work with the person.

The uncomfortable finding about interviewing is that most loops predict how good someone is at interviews.

That's not cynicism, it's what happens when four people each spend an hour forming a general impression and then average their impressions together. You've measured a single trait, four times, and the trait is legibility under pressure. Some excellent engineers have it. So do some people who will be a disaster in month five.

Fixing it isn't about finding better questions. It's about giving each stage exactly one job and refusing to let it do any others.

One job per stage

A loop where every interviewer forms a general impression is a loop where every interviewer is answering a slightly different question, and the debrief is a negotiation between four private rubrics.

So assign the jobs explicitly, in writing, before anyone talks to anyone.

The screen tests one thing: is this person plausibly in range for the level, and do they want this particular job rather than any job. Twenty-five minutes. Nothing else belongs here, and in particular technical depth doesn't, because thirty minutes of technical questioning from a recruiter or a distracted manager is a coin flip with extra steps.

The work sample tests whether they can do the actual work. Not whether they can do the actual work while being watched and narrating, which is a separate skill mostly used by consultants.

The depth interview tests the ceiling: take something they've genuinely built, and go three levels down. Why that datastore. What broke. What they'd do differently. This is where you find out whether the impressive thing on the CV was theirs.

The judgment conversation tests how they'll operate here, with the people who'll be around them. This is the one I weight highest, and it's the one that needs the most defending, so it gets its own section below.

ScreenIn range?Work sampleCan they do it?DepthCeilingJudgmentHow they operate
Each stage answers one question. The last one carries the most weight and gets defended the least.

Four stages. If you're adding a fifth, name the job it does that none of the other four cover, and if you can't, don't add it.

A work sample that respects the candidate

Ninety minutes, on a problem drawn from your actual work, with the environment they'd actually use. Their editor, their machine, the internet on, an assistant available if that's how your team works. Anything else is measuring performance under artificial constraint, and you don't ship under artificial constraint.

Scope it so that finishing is possible but unlikely. What you're reading is what they did with the time: where they started, what they chose not to do, what they checked, whether they asked a question when the requirement was ambiguous (it should be ambiguous, deliberately, in one place).

Score it against a rubric written before the first candidate saw it, with the name off the submission if you can manage it. That's not ceremony. Free-form impressions of code drift toward style preferences within about three candidates, and blind scoring is the cheapest correction available.

And pay for it if it's longer than two hours, or don't make it longer than two hours. A five-hour take-home is a tax on people with caring responsibilities and a filter for people with free evenings, which is a demographic and not a skill.

The thing a work sample will not tell you is how somebody behaves in month six, on a Tuesday, when the requirement changed and they disagree with the decision. That's the gap the next stage exists to close.

The stage I weight highest, and why it isn't culture fit

The single interview I've found most predictive is a structured conversation between the candidate and two or three of the people who'll actually work with them, about judgment: real situations from your world, no right answers, arguing them through together.

Say that in public and someone will point out, correctly, that "culture fit" interviews have a long history of producing homogeneous teams full of people the interviewer would enjoy a drink with. They're right about the failure mode. The word is genuinely bad and I'd get rid of it.

The distinction I hold is that this stage is testing values and judgment, and neither is a synonym for similarity.

Judgment: here's a situation with no clean answer. The deadline is Friday, the migration isn't safe, the customer has been promised. What do you do, what do you need to know first, who do you tell. There's no correct response and it isn't a trick. You're listening for how they reason under a real constraint, whether they consider people who aren't in the room, and whether they change their mind when someone gives them a fact they didn't have.

Values: how do they talk about a decision that went wrong. About someone they disagreed with. About work they inherited that was bad. That last one is my favourite tell, because contempt for the previous team's code is extremely common and predicts a lot of downstream friction.

What makes this stage work rather than degrade into vibes is the same structure you'd put on any other stage. The same three or four scenarios for every candidate at a level. Written scoring against named criteria, submitted before discussion. And a rule that "I liked them" and "they'd fit in here" are not accepted as evidence, ever, from anyone. If an interviewer can't point to a specific thing the candidate said, the score doesn't count.

Run it with the team rather than with management, because the team is the instrument. They're the ones who will be in code review with this person for two years, and they have a sharper read on the difference between disagreement and dysfunction than any panel of directors does.

The failure mode I watch for: a team that keeps scoring down people who are unlike them in ways that have nothing to do with judgment. That's visible in the written scores if you're looking, and it's a coaching problem for the interviewers rather than a reason to abandon the stage.

The move I use, and its limit

The one I've run for years: state a strong opinion on a live topic inside the candidate's own domain, and watch how far the pushback goes.

It works because of an asymmetry. I hold strong opinions and I hold most of them loosely, and the candidate has no way of knowing the second half. So what they're deciding, in the moment, is whether disagreeing with a senior person in an interview is worth the risk. Some rehearse the point back to me with better vocabulary. Some ask what I'm assuming. A few say they think I'm wrong and explain why, and then keep going when I push back once.

Calibrate the expectation to the authority of the role. For an engineer, a well-posed question is plenty. Hiring a VP, I want a real exchange, because the job is largely disagreeing with people who outrank them and being useful while doing it.

Now the limit, which matters more than the technique. I have this instrument. I use it. And I have still hired too many people on comfort, because comfort is a cognitive discount and running a test doesn't oblige you to act on the result. The signal arrives and the same person who designed the test discounts it.

So don't read this as the fix. A disagreement exercise produces evidence, and evidence only binds when somebody other than the decision-maker can pick it up and stop the hire. That's the mechanism. This is just the instrument that hands them something to hold.

What to do with disagreement between stages

The interesting candidates produce a split: strong work sample, mixed judgment conversation, or the reverse.

My rule is that the judgment conversation wins, because the failure modes it catches are the expensive ones. A capable engineer with poor judgment about people or risk costs you far more than a slightly slower engineer who everyone can work with, and the first kind is much harder to notice from the outside once they're in.

The exception is a split caused by nerves rather than substance. A candidate who was tense in a conversation with three strangers and excellent in the code is common, particularly for people early in their career or interviewing in a second language. Rerun the stage in a different format rather than resolving it from the scores, because you'd be resolving it against the wrong variable.

The cheapest improvement available

Write down, for the last ten people you hired, what the loop predicted and what actually happened.

Almost nobody does this, which is why interview loops improve so slowly. You'll find that one stage carried nearly all the signal and one carried none, and the one carrying none is usually the one everyone enjoys running.