Skip to main content

All posts

Hiring4 September 20267 min readBy Skillbricks Team

Technical skills assessment: what measures skill and what just filters

What a technical skills assessment should measure, four formats compared by evidence quality and candidate cost, and how to score one so it holds up.

hiringassessmentinterviewsdevops

A technical skills assessment is a structured evaluation of whether a candidate can actually do the technical work a role requires, putting evidence where the funnel would otherwise rely on claims. Done well, it is the highest-signal step in a hiring funnel. Done the way most of the market sells it, it measures something adjacent to the job and filters out some of the people you most wanted to meet.

That second sentence is the part the vendor landscape will not tell you, because most of what is sold under this label is a quiz bank or an algorithm puzzle with a leaderboard. So this guide is organised around the question that actually decides whether your assessment works: what evidence of skill does each format really produce, at what cost to the candidate, and what does that cost do to your funnel?

What should a technical assessment actually measure?

One thing: whether the candidate can do the work this role does on a normal Tuesday. Not whether they remember terminology, not whether they can invert a binary tree at speed, and not whether they had the free evenings to grind a question bank. Selection research has long placed job-relevant work samples among the more predictive assessment methods, and the recent re-analyses that revised the field's headline numbers downward have only sharpened its core lesson: instruments built from the job's real tasks and scored consistently generalise; instruments built from proxies mostly predict performance on the proxy.

For technical roles this distinction has teeth. Knowing what a liveness probe is belongs to vocabulary. Diagnosing why one is killing a healthy service, under time pressure, with incomplete information, belongs to the job. A candidate can hold the first without the second; more painfully for your funnel, plenty of strong operators hold the second while being mediocre at the first, because operating skill is built in incidents rather than flashcards. Whatever format you choose, the bar to hold it against is fidelity: how closely does the thing being scored resemble the work being hired for?

Which assessment format gives the most signal?

Four formats are what the market will mostly offer you. (Structured technical interviews and pair-programming sessions are close cousins of the last of them, and portfolio review is its asynchronous relative; the comparison below extends to those naturally.) They differ less in price than in what evidence they generate and who they silently exclude.

The four technical assessment formats plotted by evidence quality against cost to the candidate, with live work samples the only format in the high-signal, bounded-cost corner

Quiz banks and knowledge tests. Cheap, instant, easy to run at volume, and legitimately useful for one narrow purpose: confirming baseline vocabulary in high-volume screening. Their ceiling is hard, though. Answers are searchable, banks leak, and a good score proves familiarity with terminology, not the ability to apply it. Treat a quiz as a literacy check, never as evidence of competence.

Algorithmic coding tests. The industry default, and for a minority of roles a fair proxy. But for infrastructure, operations, and reliability roles the correlation runs thin: the daily work is diagnosis, systems reasoning, and careful change under pressure, none of which appears in a timed dynamic-programming exercise. These tests also carry the strongest preparation bias of any format; they reward recent grinding, which quietly filters for candidates with the most free evenings rather than the most relevant judgement. We wrote about the cost of this style of assessment here.

Take-home projects. High fidelity when the brief is genuinely job-shaped, and the format most respectful of thinking time. Two common failure modes. First, scope creep: briefs advertised as "about four hours" have a way of swallowing weekends, and the better a candidate's options, the less homework they will sit for. Second, provenance: you receive an artifact with no view of the process that produced it, and AI assistance means an artifact on its own no longer establishes who produced it. A take-home tells you what a candidate can deliver; it cannot tell you how, or how much of it was them.

Live, real-environment work samples. The candidate works a realistic task in a real environment, a broken deployment to diagnose, a service to fix, and the assessment observes the process: what they check first, how they form and discard hypotheses, what they do when the first theory is wrong. Of the four, this is the only format where the evidence includes the process, not just the artifact, which is what makes findings possible rather than opinions. "Live" by itself buys nothing, though: the value is task fidelity plus process observability plus consistent tasks and criteria across the pool, and a live session scored by vibes is theatre. Its honest costs: building realistic environments is real engineering, evaluation needs a rubric rather than a leaderboard, and a live session asks for a scheduled, bounded hour rather than an anonymous click-through.

How do assessments lose strong candidates?

Every assessment is a transaction: you are spending the candidate's time and goodwill to buy evidence. Strong candidates price that transaction sharply, because they have alternatives. Three rules keep the exchange fair and the funnel intact.

Ask for hours in proportion to commitment. A multi-hour gauntlet before any human contact reads as disrespect, and the candidates best placed to walk away from it are precisely the ones with other offers. Be transparent about what happens with the result: an assessment that produces a score the candidate never sees is extraction; one that produces feedback, or evidence the candidate keeps, is an exchange. And place the assessment where it replaces a stage rather than adding one; testing skills you then re-test in three more interviews signals that nobody trusts the instrument, and candidates notice.

An assessment is you spending the candidate's goodwill to buy evidence. Buy something the interview loop cannot, and pay a fair price.

What makes an assessment score defensible?

The instrument is only half the assessment; the other half is how you read it. Unstructured evaluation, a senior engineer eyeballing a submission and delivering a verdict, reintroduces every bias the assessment existed to remove, and it does so inconsistently across candidates, which is where both unfairness and legal exposure live. The fixes are unglamorous: a rubric written before the first candidate, equivalent tasks scored against the same criteria for everyone in the pool, scores recorded with reasons, and periodic calibration against on-the-job outcomes of past hires. Consistency is the start of defensibility, not the whole of it; job-relatedness, validation evidence, monitoring for adverse impact, and reasonable accommodations are what typically makes an instrument defensible where you hire.

Process evidence makes rubrics dramatically easier. A submitted artifact supports opinions; a record of how the candidate worked, the diagnostic path, the checks before the fix, the recovery from a wrong first theory, supports findings. If your format cannot see process, your rubric is scoring handwriting.

The honest limits of any assessment

An assessment samples a few hours of decontextualised work; the job is months of collaboration, ambiguity, and compounding context, and no instrument closes that gap entirely. False negatives are real: some excellent engineers perform badly under observation, and a funnel tuned only for precision will quietly discard them. Fidelity also has a price, and for a high-volume, low-complexity role, a modest instrument honestly applied beats an elaborate one you cannot afford to run consistently. The buying checklist beyond fidelity is boring and decisive: evaluator time per candidate, consistency at your volume, accessibility, and what evidence the candidate gets back. The discipline is matching the instrument's cost to the decision's stakes, then weighing what it says alongside the rest of your loop and checking its predictions against how hires actually perform.

Where we fit in

Everything above is the checklist we built SkillBricks against. Our assessments are live work samples in real environments (infrastructure roles are the wedge), with the process observed and scored against a rubric, so the evidence is findings rather than opinions. And because candidates keep what they earn, verified skills on a wall that speaks before their name does, the transaction is an exchange: your funnel starts from people whose skills are already evidenced, and their effort was not spent solely on your pipeline. The same reasoning in role-specific detail, including when a simpler instrument is genuinely the right call, is in our guide to assessing DevOps engineers; the wider case for skills-led funnels is in the skills-based hiring piece. If your current assessment is a quiz bank with a leaderboard, the question worth sitting with is the one this page started from: what is it actually measuring?