High-volume roles: screen everyone, not just the top of the pile
What actually happens to two hundred applications, why the shortlist owes so much to formatting, and how a written screen changes who gets looked at.
Two hundred applications. Time for twenty calls.
Everyone knows what happens next, and almost nobody says it out loud. There is no reading of two hundred applications. There is a fast pass, a few minutes a head at best, and a shortlist that owes as much to formatting, familiar company names and the order the applications arrived in as it does to what any of these people can actually do.
That is not a criticism of the recruiter. It is arithmetic. The pile grew and the day did not.
This page is about the move that changes that arithmetic: making the screening step cheap enough that you can apply it to the whole pile instead of to the part you had time for.
What actually happens to the pile
Ask any recruiter how long they spend on an application in a high-volume search and the honest answer is somewhere between a few seconds and a couple of minutes. Nobody defends that number. It is what fits.
And in those seconds, the things that are quick to see win. Company names you recognise. A layout that scans. A title that matches the one in the job description. A gap explained in the first line rather than the fourth.
Every one of those is weakly related to whether the person can do the job, and strongly related to how much time and coaching went into the document. That would be a bad trade in any market. It is a worse one now that the document can be generated to order.
The reading budget, with the numbers made explicit
Say 200 applications arrive over three weeks. Give each one a genuinely fair read: skim the CV, open the portfolio link, form a view. Call it six minutes. That is 20 hours of work, which is most of a working week that the recruiter does not have, on top of everything else on their desk.
So the real budget is closer to 30 seconds each, or under two hours total. At that speed you are not evaluating people, you are sorting documents by how quickly they announce themselves.
Now look at the other end of the same process. Say 20 of those people get a call, 8 reach an interview, and 1 gets hired. The 180 who were never called were filtered by two hours of skimming. Whether the best candidate was among them is not knowable, because nobody looked.
The uncomfortable part of that calculation is not the hours. It is the last sentence. The cost of a fast pass is not measured in time saved, it is measured in candidates you never evaluated and will never know you missed.
Why the pile got bigger
Applying became close to free, and both sides of the market noticed.
ZipRecruiter's 2026 AI Employer Report found that just under half of surveyed employers say AI has increased application volume per role. The same research found that a quarter of employers believe they can almost always tell when AI was used in an application, with most of the rest saying they can sometimes tell, which is a polite way of saying the signal is now unreliable in both directions.
Two consequences follow, and they pull against each other.
The pile is larger, so the fast pass gets faster. And the documents in the pile are more uniform, so the fast pass has less to go on. The filter is being asked to do more work with worse inputs, which is exactly when a process starts producing results that look reasonable and are close to random.
The move: screen the whole pile
The instinct is to filter harder before screening. More keywords, stricter requirements, an experience floor. That reduces the pile, but it reduces it along the same axis that was already unreliable.
The other move is to make the screening step cheap enough to apply to everyone. Not a shortlist of forty. All two hundred, including the ones whose CV you would have discarded on formatting.
That sounds expensive until you notice what the cost actually is. Writing the questions is a one-time job per role, not per candidate. Sending the link is one action for the whole list. What scales with the pile is the reading of the answers, and that is the part where a score and a per-competency breakdown do real work: they give you a defensible order to read in, rather than a random one.
The point is coverage, not speed. Everyone answered the same questions. The person at position 180 in your inbox got exactly the same chance to show what they know as the person at position 3.
See what a written screen looks like
A new account opens with a filled-in example questionnaire and five candidates, so you can read a real report before writing a single question of your own.
Start freeWhat to ask when the pile is large
The temptation in high volume is to ask more, because there are more people to separate. Do the opposite.
Three or four questions. Say in the invitation roughly how long it should take, and mean it. A questionnaire that takes an hour will be abandoned by exactly the candidates who have other options, and you will have filtered for free time rather than for ability.
What works at volume:
- One question that has a wrong answer. A small scenario from the actual job with a decision in it. "A customer says the export is broken. It is not broken, they are looking at the wrong report. What do you write back?"
- One question about something they did. Not a summary of the CV. A single specific instance, with what went wrong in it. Specifics are the hardest thing to produce without the experience behind them.
- One question that reveals the trade they would make. "You have time to do one of these two things properly or both of them badly. Which, and why?"
- Optionally, one screening constraint in writing. Notice period, working hours, location requirements, whatever the actual gate is. Getting it in writing early saves a call whose only purpose was to discover a mismatch.
Notice that none of these can be answered by restating a job description, and none of them require the candidate to write an essay.
Yes, some of them will use a model
At two hundred applications, this objection arrives immediately, and it deserves a straight answer rather than a reassurance.
Some candidates will paste your questions into a chatbot. More of them will use one to tidy up an answer they wrote themselves, which is a different thing and not obviously bad. Stack Overflow's 2025 Developer Survey found that 84% of respondents use or plan to use AI tools while 46% distrust the accuracy of the output, and although that survey is about developers, the pattern is now general: people use the tool and then check it. A candidate who does that is behaving the way they will behave at work.
So the question is not whether the tool was used. It is whether the person can back up what came back.
That is what Eneya reads for. Either the answers name something you could check, a company, a product, a project with details that only make sense there, or they do not. The report shows that as a High or Medium reading next to the score, with the fragments that gave it, and it does not try to guess whether a model was involved.
That reading is a reason to ask one more question, not a reason to reject. Medium means nothing named could be verified, and at volume you will meet plenty of honest answers like that. The follow-up is cheap and it is the same one every time: ask them to walk through the specific case they described. Anyone who lived it can. Anyone who did not will move to generalities within two sentences.
There is also a design point hiding here. Questions that ask what the candidate did, on what system, and what went wrong are much harder to answer from a prompt than questions that ask what they think about a topic. The defence against generated answers is mostly in the question, not in the detector.
What the number is, and what it is not
If a score is going to decide reading order for two hundred people, it is worth knowing exactly what produced it.
The model reads each answer against the competencies you defined and reports what it found, competency by competency, with the part of the answer it based that on. The final number is then computed from those assessments by fixed code in the product, not by the model, which means the same assessments always produce the same number, and the report shows the breakdown that led to it. Each stored evaluation also carries the version of the rule it was scored with, so a number from three months ago can still be explained rather than guessed at.
Two honest limits come with that.
The first is that the top of the range compresses. Two strong candidates can land within a point or two of each other, and the difference between them is not in the number, it is in the text and in the follow-up questions the report suggests. Use the score to decide who to read, not to decide between the people you have read.
The second is that a score can only reflect what the questions asked. If the questionnaire covers three competencies and the role needs five, the number is a confident answer to a question you did not mean to ask. When coverage is too thin to support a number, Eneya says so instead of producing one, which is occasionally annoying and always better than the alternative.
Fairness stops being an aspiration
There is a second reason to do this that has nothing to do with time.
At the point when someone asks why one candidate progressed and another did not, a written screen gives you an answer that exists. Everyone was asked the same questions. The criteria were fixed before the answers arrived. The answers themselves are still there to read.
That is the same mechanism that the selection research keeps pointing at. In the reanalysis by Sackett and colleagues published in the Journal of Applied Psychology in 2022, structured interviews came out with the highest mean validity among commonly used selection procedures, after correcting a statistical error that had inflated earlier estimates for decades. The finding is about interviews rather than questionnaires, and the numbers are averages across many studies rather than a promise about your next hire. What transfers is the structure: same questions, criteria set in advance, a record of what was asked.
Compare that to defending a shortlist built by a two-hour skim. Not because the skim was unfair in intent, but because there is nothing to point at afterwards except a recollection.
One more thing about ordering. The default sort in a candidate list is usually recency, which answers the question "what arrived last" rather than "who needs me next". At two hundred applications that difference is not cosmetic: the person who applied on day one is at the bottom of the screen by day three, and the fast pass compounds the accident. A list ordered by what needs a decision, with the score visible next to each row and the ones still being evaluated marked as such, is a different working surface. It is the difference between a queue and an inbox.
The candidate side of a large pile
High-volume hiring has a reputation among candidates, and it is deserved. Most of the 180 people who were never called also never heard anything.
Bullhorn's 2026 GRID Industry Trends Report puts speed and responsiveness at the centre of what candidates value, and quotes one agency leader describing the shift from a black hole to a response on something the candidate spent real time on. That is the bar now, and it is not a high one.
A written screen changes the interaction in a small but real way. The candidate does something concrete, on their own time, on any device, without recording themselves or installing anything. They are being asked about the work rather than being asked to wait.
It does not make a rejection pleasant. It does make the process one where the person had a chance to show something, which is more than a CV in a queue offers them.
There is a version of this that goes wrong, and it is worth naming. If the questionnaire is long, the questions are generic, and the invitation says nothing about why it exists, then you have added a chore to the front of your funnel and the people who skip it will be disproportionately the ones with choices. Everything about high-volume screening depends on the step being small enough to be worth doing for a job the candidate is not yet sure they want.
What this does not do
It does not replace the interview. It changes who gets one.
It does not eliminate drop-off. Some people will not answer, and the number will be larger than you expect the first time. Keep the questionnaire short, say why it exists, and give a real deadline. Treat drop-off as feedback on the length of your questionnaire before treating it as feedback on the candidates.
It does not read the answers for you. A score puts the pile in an order that is defensible. The evidence is the text, and at least the top of the list has to actually be read by a person, or you have built a faster way to make the same mistake.
It does not scale for free. You pay per evaluation, which means the cost of screening two hundred people is a real number rather than a rounding error. That is a deliberate part of the design: a per-seat subscription would make the marginal candidate free and quietly encourage sending the thing to everyone forever. See pricing for what that looks like, including the free tier, which is enough to run one role end to end before deciding anything.
Where this breaks down
Roles where written expression is not job-relevant. Warehouse, kitchen, field service, shift work. Asking people to write when the job never asks them to write is a filter for something other than the job, and at volume that filter will systematically remove good candidates. If you use written screening for these roles, keep the questions concrete and operational, and weigh the answers accordingly.
Language. A written screen in a second language measures the second language too. Sometimes that is legitimate, because the job runs in that language. Often it is a requirement nobody examined. Decide which one it is before you send it, not after you look at the scores.
Accessibility. Some candidates need more time or a different format. Say in the invitation that they can ask, and make sure someone answers when they do.
Very short-cycle roles. If you need someone starting Monday, a step that takes a day of candidate time is not free. In that case a written screen is better used for the pipeline you are building for next month than for the fire you are putting out this week.
Starting this week
- Take the role with the worst ratio of applications to interview slots. That is where the fast pass is doing the most damage.
- Write three questions. Half an hour, once, reused for every candidate on that role and the next time you open it.
- Send it to everyone, including the applications you would have discarded after ten seconds. This is the whole point, and it is also where the surprises come from.
- Read the answers from the top of the scored list, but read at least a few from the middle. That is how you find out whether the order is any good.
- When you make your calls, write down why before you look at the number.
Two hundred applications will still be two hundred applications. The difference is that all of them got asked, the shortlist has something behind it, and the person who writes plainly but formats badly is no longer invisible.
If the format itself is the question, the longer argument for written answers over recorded ones is in written pre-screening or one-way video. If you are hiring engineers specifically, the same idea applied to protecting senior interview time is on the engineering hiring page.