In This Article 9 min read
Key Takeaways
We scored 1,848 recorded customer service answers and 699 timed written assessments from 670 applicants. The written test produced a 20.9 point spread between two groups of candidates. On a recorded call, that spread shrank to 3.7 points.
Almost every offshore customer service hire starts the same way: an unproctored written assessment, taken at home, on the candidate’s own machine. It is cheap, it scales, and it produces a tidy number you can sort a spreadsheet by.
Our screening platform records something most assessment tools do not keep: what happened in the browser while the candidate was answering. Tab switches. Paste events. Time on task. Connection speed. When we lined that telemetry up against the same candidates’ scored voice recordings, the written scores stopped looking like a measure of ability.
This report is the analysis. All of the data is first-party, from applications submitted directly to Armasourcing between 2026-06-05 and 2026-07-27. The method and its limits are set out before the findings, not after.
Method, and what this data cannot tell you
Every candidate in this dataset applied for a remote customer service or customer-service-adjacent role and completed two things: a timed written assessment taken in the browser, and a set of recorded voice answers to scenario prompts. Voice answers are transcribed and scored against a five part rubric covering fluency, clarity, grammar, tone and accuracy, each out of 20.
- 1,848 scored voice recordings from 670 candidates, an average of 2.8 per person.
- 699 timed written assessments with full browser telemetry.
- 673 candidates completed both and form the paired set behind the headline finding.
- Window: 2026-06-05 to 2026-07-27.
Limits you should hold this against:
- It is close to one employer. 1,635 of the 1,848 recordings (88 percent) came from a single hiring push for a United States moving and storage company. This is not a representative sample of the Philippine customer service labour market, which the IT and Business Process Association of the Philippines puts well over a million workers. It is a deep look at one real funnel.
- The window is short. Roughly seven weeks, in one hiring season.
- Scores come from a rubric applied by a language model, not from human raters. We ran an independent blind second scoring pass over a stratified 40 answer sample to check it. Agreement was moderate: correlation of 0.357, mean absolute difference of 9.8 points out of 100, and a Cohen’s kappa of 0.25 on the binary question of whether an answer clears 70, which is fair rather than strong. Group comparisons of the size reported below are robust to that much noise. Individual scores should not be read as precise, and we do not report any finding here that depends on individual precision.
- We removed 54 recordings before analysis. 16 never finished scoring, and 38 returned an identical value on all five rubric dimensions, which is a scoring fallback rather than an assessment. A further 52 recordings (2.8 percent of those kept) sit at exactly 10 on fluency, clarity and grammar together, which we suspect is partly the same artefact. We kept them, because some answers genuinely earn that, and we flag them here so you can discount accordingly.
- A paste event is not proof of anything. Candidates paste prepared notes, their own names, portfolio links and rehearsed answers. We report the behaviour and what it correlates with. We do not call it cheating, and neither should anyone citing this. The APA’s guidance on the rights and responsibilities of test takers is a reasonable starting point for anyone drawing conclusions about individuals from assessment telemetry.
If you are hiring one or two people rather than a team, the same mechanics apply, and our page on Filipino virtual assistants for customer service covers that end of the market.
The finding: the written test’s ranking does not survive a recorded call
Split the 673 paired candidates by whether the browser recorded a paste event during the written assessment. Then look at how each group scored on the written test, and how the same people scored on their recorded voice answers.
| Group | n | Written assessment | Recorded voice call |
|---|---|---|---|
| Pasted at least once | 252 | 69.0 | 74.6 |
| No paste detected | 421 | 48.1 | 70.9 |
| Gap | 20.9 | 3.7 |
On the written assessment, candidates who pasted outscored candidates who did not by 20.9 points. That is an enormous spread. It is the difference between the top of a shortlist and the reject pile.
Put the same people on a recorded call and the gap falls to 3.7 points. 82.3 percent of the advantage disappears.
The residual 3.7 point difference is statistically detectable, but it is smaller than the 9.8 point margin between our own two scoring passes. We would not build a hiring decision on it, and we are not asking you to.
We checked whether this was an artefact of how we cleaned the data or which roles we included. It is not. Across five specifications, from the raw data through to the single largest role with every suspect score removed, the collapse lands between 82 and 85 percent.
| Specification | Candidates | Written gap | Voice gap | Gap closed |
|---|---|---|---|---|
| Raw, all roles | 678 | 20.8 | 3.7 | 82.1% |
| Fallback scores removed | 673 | 20.9 | 3.7 | 82.2% |
| Fallback plus suspect cluster removed | 667 | 21.4 | 3.9 | 82.0% |
| Largest single role only, raw | 561 | 21.0 | 3.2 | 84.7% |
| Largest single role only, fully cleaned | 556 | 21.4 | 3.3 | 84.6% |
How common is this? More common than most hiring managers assume
Across all 699 written assessments:
- 36.6 percent recorded at least one paste event.
- 42.9 percent switched away from the assessment tab at least once.
- 21.2 percent switched tabs five or more times. The highest single sitting recorded 54 switches.
Tab switching behaves like a dose. A candidate leaving the tab once or twice scores about the same as one who never leaves. Past five switches, the score climbs sharply.
| Tab switches | n | Mean written score |
|---|---|---|
| 0 | 399 | 52.9 |
| 1-4 | 152 | 53.5 |
| 5-9 | 69 | 61.6 |
| 10+ | 79 | 68.7 |
One detail complicates the lazy reading of this. Candidates who pasted spent longer on the assessment, not less: 28.9 minutes on average against 22.7 for candidates who did not. Whatever is happening, it is not people rushing through. It looks like effort, applied to the tool rather than to the answer.
What the written test actually spreads people across
The written scores look healthy in aggregate. The mean is 55.7, the median 60.0, and 36.2 percent of submissions clear 70. The distribution is wide, which is what you want from a screening instrument.
| Score band | n | Share |
|---|---|---|
| Under 40 | 235 | 33.6% |
| 40 to 54 | 85 | 12.2% |
| 55 to 69 | 126 | 18.0% |
| 70 to 84 | 175 | 25.0% |
| 85 or more | 78 | 11.2% |
The problem is not that the spread is narrow. It is that a large part of it tracks browser behaviour rather than the candidate. A test that separates people cleanly is only useful if it separates them on something you meant to measure.
What a recorded call measures that a written test cannot
Voice answers are scored on five dimensions. Averaged across 1,848 recordings, they do not come out level.
| Dimension | Mean out of 20 | Share of maximum |
|---|---|---|
| Tone | 16.0 | 80.0% |
| Accuracy | 14.7 | 73.6% |
| Clarity | 14.6 | 73.2% |
| Fluency | 14.1 | 70.4% |
| Grammar | 13.7 | 68.6% |
Tone is the strongest dimension at 80.0 percent, and grammar the weakest at 68.6 percent. The pattern is consistent: warmth, patience and the instinct to reassure an anxious customer show up reliably. Article use, tense agreement and preposition choice are where points go. In CEFR terms that is the signature of a strong B2 population, comfortable and effective in real conversation while still making visible grammatical slips.
That ordering matters for anyone building a screening process, because it is the reverse of what a written test is good at. Written assessments measure grammar well and tone barely at all. They are strongest exactly where this candidate pool is weakest, and blind exactly where it is strongest. If you want to know more about which numbers actually predict customer service performance, our guide to contact center KPIs and metrics covers the operational side, and in-house versus outsourced contact centers covers where the screening burden lands.
The constraint nobody screens for
Voice work has a hardware floor that written work does not. We measure connection speed in the browser at assessment time.
| Download speed | n | Share |
|---|---|---|
| Under 25 Mbps | 230 | 33.3% |
| 25 to 49 | 271 | 39.2% |
| 50 to 99 | 174 | 25.2% |
| 100 or more | 16 | 2.3% |
Median download speed is 36.4 Mbps and median latency 68 ms, which is comfortable for voice. But 33.3 percent measured under 25 Mbps, and only 2.3 percent reached 100 or more. Average typing speed was 43.3 words per minute, which matters for anyone expected to document a call while taking the next one.
A candidate can be excellent and still be unable to hold a clean call from where they live. That is a solvable problem, but only if you measure it before you hire rather than after. Connectivity is one of the practical reasons teams cluster in particular cities, which we cover in our guide to call center outsourcing in the Philippines and on our Philippines contact center page.
What we would take from this
We are a staffing company, so treat the following as interested advice and check it against your own funnel.
- Assume an unproctored written score is partly a measure of tool access. Not entirely, and not for every candidate. But a 20.9 point spread that closes to 3.7 on a recorded call is not a small correction. The International Test Commission’s guidelines on computer-based and internet-delivered testing have warned about exactly this control gap for years, and they predate widely available generative tools.
- Record telemetry even if you do nothing with it. Paste events, tab switches and time on task cost nothing to capture and change how you read a score.
- Make candidates speak, early. A short recorded answer to a real scenario is the cheapest high-signal step available for a customer service role, and it is the step this data says survives contact with reality. This matters most for inbound voice work, where the first thirty seconds of a call set the outcome.
- Do not screen on self-declared English. We looked hard at whether self-described fluency predicted measured performance and could not separate the groups. Our scoring is not precise enough for us to publish that as a finding, which is itself the point: if we cannot measure the difference at this sample size, an application form dropdown certainly cannot.
- Test the connection before the offer, not after. A third of this pool would need an upgrade to run voice reliably.
Reuse and citation
All figures in this report may be reproduced with attribution and a link to this page. If you are citing it, this line works:
Armasourcing, The Customer Service Screening Report 2026, based on 1,848 scored voice recordings and 699 written assessments from 670 customer service applicants.
Underlying recordings and transcripts are not published. Every figure here is an aggregate, and no candidate is identifiable from anything in this report.
Related first-party research: our Filipino Virtual Assistant Salary Report covers what candidates in this market ask to be paid. If you are working out how to structure a customer service team offshore, start with the contact center outsourcing hub, read how to choose a contact center outsourcing company, or compare building in-house against outsourcing.
Frequently asked questions
How many candidates does this report cover?
670 customer service applicants, who between them produced 1,848 scored voice recordings and 699 timed written assessments between 2026-06-05 and 2026-07-27. The paired analysis behind the headline finding covers the 673 candidates who completed both.
What share of candidates pasted into their assessment?
36.6 percent of the 699 written assessments recorded at least one paste event, and 42.9 percent recorded at least one tab switch. Pasting is not proof of misconduct. Candidates paste prepared notes and rehearsed answers as well as anything else.
Does a written assessment predict customer service performance?
In this data it predicts less than it appears to. Candidates who pasted scored 20.9 points higher on the written test than those who did not, but only 3.7 points higher on a recorded voice call. 82.3 percent of the written advantage did not survive the change of format.
Where are Filipino customer service candidates strongest and weakest?
Across 1,848 scored recordings, tone is the strongest dimension at 80.0 percent of maximum and grammar the weakest at 68.6 percent. Warmth and customer handling instinct score consistently higher than grammatical precision.
Is this data representative of the whole Philippine customer service market?
No. 88 percent of the recordings come from a single hiring push for one United States moving and storage company, over roughly seven weeks. It is a deep look at one real funnel rather than a market-wide survey, and it should be cited that way.
Need a VA who already understands your industry?
We don’t place generalists. Our VAs are matched and trained for the specific workflows of your sector.
See industry VAs →




