🇸🇬 HireDeveloper.sg

We Dropped the 8-Hour Take-Home for a 90-Minute Paired Trial — Acceptance Went from 54 % to 81 %

Two engineers reviewing TypeScript code together on a shared screen
William

William

Talent Sourcing Expert · 3 September 2026 · 15 min read

TL;DR

  • • Same 48 candidates, two formats. Take-home: 22 submitted, 4 hired, 54 % acceptance. Paired trial: 42 completed, 9 hired, 81 % acceptance.
  • • A take-home measures the artefact. A trial measures the engineer. You are hiring the engineer.
  • • You cannot know who wrote a take-home — and since 2025 that is decisive, not pedantic.
  • • Build one deliberate ambiguity into the task. What a candidate does with an unclear requirement is the single most informative moment.
  • • Brief the interviewer: answer every domain question, never answer a design question.
  • • Anything over ~2 hours of a candidate’s time should be paid. The unpaid take-home filtered out exactly the employed, in-demand people we wanted.

We ran both formats against the same 48 shortlisted TypeScript candidates over eleven months. The take-home lost 54 % of them before submission. The 90-minute paired trial lost 12 %. But the conversion gap is not the point — the two formats measure different things, and only one measures what you are hiring for.

We ran both formats against the same shortlist of 48 TypeScript candidates for Singapore-based roles over roughly eleven months. The take-home lost 54 % of candidates before submission and converted 4 hires. The 90-minute paired trial lost 12 %, and converted 9.

The interesting part is not the conversion difference. It is that the two formats measure different things, and only one of them measures the thing you are hiring for.

Step 1 — Write down what you cannot learn from the CV and the first call

Every evaluation format is expensive, so it should have a job. Before designing anything, name the two or three things you genuinely cannot assess after a technical conversation.

For senior TypeScript roles these are consistent: does this person choose sensible boundaries between modules when the requirement is ambiguous, do they reach for types as a design tool or as a formality, and what do they do when something does not work?

Notice that none of those are answered by a finished, polished artefact. They are all questions about process, and process is exactly what a take-home discards.

Step 2 — Understand what each format can physically observe

A take-home produces a finished submission. You can assess final code quality well, and almost nothing else. You cannot see which of two designs the candidate rejected, how long they were stuck, what they looked up, or whether the requirement confused them.

You also cannot know who wrote it. This has always been true and is now decisive, because a competent engineer with an assistant produces indistinguishable output from a strong engineer working alone. If your take-home was calibrated before 2025, it is measuring something different now than it was then.

A paired trial produces a recording of decisions. You see the choice being made, the moment of being stuck and what happens next, and the response when you change a requirement halfway through. It is weaker on final polish and stronger on everything you actually listed in step 1.

WHAT EACH FORMAT CAN ACTUALLY OBSERVESIGNALTAKE-HOMEPAIRED TRIALFinal code qualityStrongPartialHow they choose between two designsInvisibleDirectly observedBehaviour when stuckInvisibleDirectly observedResponse to a changed requirementInvisibleDirectly observedWho actually wrote itUnknowableCertainThe take-home measures the artefact. The trial measures the engineer. You are hiring the engineer.

Step 3 — Design a 90-minute trial with a real, small, deliberately ambiguous task

Ninety minutes is the number we settled on. Sixty is too short for anything but a puzzle; two hours pushes into unpaid-work territory and the drop-off rises.

The task should come from your actual codebase, reduced. Not a puzzle, not an algorithm. For TypeScript roles we use something like: here is a module with a loose type surface and two consumers, extend it to support a third case. It takes us ten minutes to prepare and it is recognisably the job.

Build in one deliberate ambiguity. A requirement that can be read two ways. What the candidate does with it is the single most informative moment in the session — strong candidates notice and ask; weaker ones pick one reading silently and build on it.

Step 4 — Decide who pairs, and how they behave

The person pairing should be someone the candidate would actually work with, and they need a brief, because the default failure mode is an interviewer who either takes over or goes silent.

The rule we give: answer any question about the domain immediately and completely; never answer a question about the design. If the candidate asks what this field represents, tell them. If they ask whether they should split the module, ask what they are weighing.

Say at the start, explicitly, that asking questions is expected and not penalised. Without that sentence, roughly half of candidates will silently guess rather than ask, and you will have measured their assumption about interview etiquette rather than their engineering.

Want a trial format that candidates actually complete?

We design and run structured technical trials for Singapore teams, then hand you the evidence — decisions observed, not artefacts submitted.

Start now — see how we vet

Step 5 — Score during the session, on four axes, before you discuss

Scoring afterwards from memory produces a single overall impression dressed up as analysis. Score live, on a shared sheet, before any conversation between interviewers.

Four axes cover it. Problem framing: did they clarify the ambiguity before building? Design judgement: were the boundaries defensible, and could they say why? Recovery: when stuck, did they narrow the problem or thrash? Communication under load: could you follow their reasoning while they worked?

One to four on each, no half points, no overall score. Forcing a number per axis prevents the common failure where a likeable candidate scores well on everything and an abrupt one scores badly on everything.

SAME 48 SHORTLISTED CANDIDATES — TWO EVALUATION FORMATSTAKE-HOME (6–8 HOURS)Invited: 48Started: 31Submitted: 22Offered: 7Accepted: 4Acceptance rate: 54 %54 % drop-off before submissionPAIRED TRIAL (90 MINUTES)Invited: 48Scheduled: 44Completed: 42Offered: 11Accepted: 9Acceptance rate: 81 %12 % drop-off before completion

Step 6 — Debrief within thirty minutes, and separate signal from preference

Debrief while it is fresh, and start by having each interviewer read their axis scores before anyone discusses. The first person to speak in an unstructured debrief sets the outcome; reading scores first prevents that.

Then apply one filter to every negative comment: is this a capability observation or a style preference? “Reached for a class where a function would do” is usually preference. “Could not explain why they split it that way” is capability. The distinction sounds obvious in writing and is routinely missed in the room.

Write the decision as a sentence with a reason. Not a score, not a recommendation — a sentence someone can disagree with in six months when the hire is going well or badly.

Step 7 — Tell them the same day, and pay for anything longer

Same-day feedback is the cheapest reputational investment available in a small market. Singapore’s senior TypeScript community is not large; candidates talk, and a team known for fast, specific responses gets easier access to the next candidate.

And a boundary worth stating clearly: if your format asks for more than about two hours of a candidate’s own time, pay for it. The 6–8 hour unpaid take-home is the reason 54 % of our invited candidates never submitted, and the ones who declined skewed strongly toward those already employed and in demand — which is to say, the ones you wanted.

If you are building out the surrounding process, our guidance on hiring TypeScript developers in Singapore covers the stages either side of the trial. Employers comparing evaluation standards across regional hubs will find parallel material at HireDeveloper.ae and JapanDev.

Calibrate the trial before you run it on candidates

An evaluation format that has never been tested against known outcomes is an opinion with a scoring sheet attached. Before running the trial on external candidates, put three of your own engineers through it: one you consider strong, one solidly average, and one who has recently struggled.

The result is usually uncomfortable and always informative. If your strongest engineer scores mid-range on design judgement, either the task is badly constructed or your assessment of them was. Both are worth discovering before you start rejecting applicants on that basis.

Calibrate the clock as well. If your own team clears the task in twenty minutes it is too easy; if they need more than seventy, the module is too large or too specific to your codebase. The target is roughly forty-five to sixty minutes of actual work inside the ninety-minute session, leaving room for the ambiguity conversation and a changed requirement.

Re-run this calibration annually, or whenever the codebase changes materially. A trial task drawn from a part of the system you have since replaced quietly starts measuring something other than what you intended.

The trial tells you about your codebase too

A side effect that few teams anticipate: the paired trial is a mirror. You put real code from your system in front of a sequence of capable outsiders and then listen to independent, unfiltered assessments of it.

After roughly ten sessions we noticed a pattern that had never surfaced internally — almost every candidate flagged the same aspect of our error-handling approach as confusing, while inside the team it had long been accepted as simply how things were. That information alone justified the format change.

Use this deliberately. Alongside the candidate scores, keep a running note of recurring criticisms of the task code itself. Anything three out of five experienced engineers independently object to is very likely a real problem rather than a matter of taste.

One fairness caveat: if you improve the task code in response to that feedback, you have changed the assessment. Record which version was used from which date, so that scores remain comparable over time. Without that note, a year later you will be comparing candidates who solved materially different problems.

When a take-home is still the right answer

Two cases. If the role is genuinely asynchronous and the work product is long-form — a technical writing component, an architecture proposal — then a written artefact is the job, and asking for one is honest.

And for candidates who explicitly prefer it. Some people interview badly under observation and produce excellent work alone. Offering both formats and letting the candidate choose costs you a little consistency and buys you access to a population that a mandatory paired trial quietly filters out.

Frequently asked questions

Is a paired trial better than a take-home for TypeScript hiring?

For most roles, yes, because the two formats measure different things. A take-home produces a finished artefact, so it assesses final code quality well and almost nothing else: you cannot see which design the candidate rejected, how long they were stuck, or whether the requirement confused them. A paired trial produces a record of decisions, which is what senior hiring questions actually turn on. Across the same shortlist of 48 candidates for Singapore roles, our 6–8 hour take-home lost 54 % of candidates before submission and converted four hires, while a 90-minute paired trial lost 12 % and converted nine. There remain two cases where a take-home is right: genuinely asynchronous roles with long-form work products, and candidates who explicitly prefer to work unobserved.

How long should a technical trial be?

Ninety minutes works well and is the length we settled on. Sixty minutes is too short for anything but a puzzle, and puzzles measure puzzle-solving rather than engineering. Two hours or more pushes into unpaid-work territory, and the drop-off rate rises sharply among exactly the candidates you most want, namely those already employed and in demand. The task itself should come from your real codebase, reduced to something that takes about ten minutes to prepare: not an algorithm, but something recognisably the job, such as extending a module with a loose type surface to support a third consumer. Build in one deliberate ambiguity, because what a candidate does with an unclear requirement is the most informative moment available.

How do you stop the interviewer from ruining a paired trial?

Give them an explicit brief, because the default failure modes are taking over the keyboard or going completely silent. The rule that works is simple: answer any question about the domain immediately and completely, and never answer a question about the design. If the candidate asks what a field represents, tell them. If they ask whether to split a module, ask what they are weighing. It also matters to say aloud at the start that asking questions is expected and will not be penalised; without that sentence, roughly half of candidates silently guess rather than ask, and you end up measuring their assumptions about interview etiquette rather than their engineering judgement.

Should candidates be paid for technical evaluation time?

Anything beyond roughly two hours of a candidate’s own time should be paid. The unpaid six-to-eight hour take-home is the direct cause of the 54 % drop-off we measured before submission, and the drop-off is not random: it skews heavily toward candidates who are already employed, in demand and have alternatives, which is precisely the population a senior search is trying to reach. A 90-minute paired trial sits below the threshold where payment becomes an issue, which is part of why completion rates are so much higher. Same-day feedback is the other cheap investment, particularly in a market as small as Singapore, where candidates talk to each other.

Rebuilding your technical evaluation in Singapore?

We design and run the trial, score it live on four axes, and hand you a decision with a reason attached — not a gut feeling dressed up as a score.

Start now — see vetted candidates