The Customer Service Screening Report 2026

Slope chart cover: 82 percent of the written-test advantage disappears when customer service candidates are put on a recorded call
In This Article 9 min read

    Key Takeaways

      We scored 1,848 recorded customer service answers and 699 timed written assessments from 670 applicants. The written test produced a 20.9 point spread between two groups of candidates. On a recorded call, that spread shrank to 3.7 points.

      Almost every offshore customer service hire starts the same way: an unproctored written assessment, taken at home, on the candidate’s own machine. It is cheap, it scales, and it produces a tidy number you can sort a spreadsheet by.

      Our screening platform records something most assessment tools do not keep: what happened in the browser while the candidate was answering. Tab switches. Paste events. Time on task. Connection speed. When we lined that telemetry up against the same candidates’ scored voice recordings, the written scores stopped looking like a measure of ability.

      This report is the analysis. All of the data is first-party, from applications submitted directly to Armasourcing between 2026-06-05 and 2026-07-27. The method and its limits are set out before the findings, not after.

      Method, and what this data cannot tell you

      Every candidate in this dataset applied for a remote customer service or customer-service-adjacent role and completed two things: a timed written assessment taken in the browser, and a set of recorded voice answers to scenario prompts. Voice answers are transcribed and scored against a five part rubric covering fluency, clarity, grammar, tone and accuracy, each out of 20.

      • 1,848 scored voice recordings from 670 candidates, an average of 2.8 per person.
      • 699 timed written assessments with full browser telemetry.
      • 673 candidates completed both and form the paired set behind the headline finding.
      • Window: 2026-06-05 to 2026-07-27.

      Limits you should hold this against:

      • It is close to one employer. 1,635 of the 1,848 recordings (88 percent) came from a single hiring push for a United States moving and storage company. This is not a representative sample of the Philippine customer service labour market, which the IT and Business Process Association of the Philippines puts well over a million workers. It is a deep look at one real funnel.
      • The window is short. Roughly seven weeks, in one hiring season.
      • Scores come from a rubric applied by a language model, not from human raters. We ran an independent blind second scoring pass over a stratified 40 answer sample to check it. Agreement was moderate: correlation of 0.357, mean absolute difference of 9.8 points out of 100, and a Cohen’s kappa of 0.25 on the binary question of whether an answer clears 70, which is fair rather than strong. Group comparisons of the size reported below are robust to that much noise. Individual scores should not be read as precise, and we do not report any finding here that depends on individual precision.
      • We removed 54 recordings before analysis. 16 never finished scoring, and 38 returned an identical value on all five rubric dimensions, which is a scoring fallback rather than an assessment. A further 52 recordings (2.8 percent of those kept) sit at exactly 10 on fluency, clarity and grammar together, which we suspect is partly the same artefact. We kept them, because some answers genuinely earn that, and we flag them here so you can discount accordingly.
      • A paste event is not proof of anything. Candidates paste prepared notes, their own names, portfolio links and rehearsed answers. We report the behaviour and what it correlates with. We do not call it cheating, and neither should anyone citing this. The APA’s guidance on the rights and responsibilities of test takers is a reasonable starting point for anyone drawing conclusions about individuals from assessment telemetry.

      If you are hiring one or two people rather than a team, the same mechanics apply, and our page on Filipino virtual assistants for customer service covers that end of the market.

      The finding: the written test’s ranking does not survive a recorded call

      Split the 673 paired candidates by whether the browser recorded a paste event during the written assessment. Then look at how each group scored on the written test, and how the same people scored on their recorded voice answers.

      The written-test gap does not survive a recorded callWritten assessmentRecorded voice callunproctored, at homesame candidatesPasted at least oncen=25269.074.6No paste detectedn=42148.170.920.9 pts3.7 ptsThe gap closes by 82.3%Mean scores out of 100. The same 673 candidates, measured both ways.
      Mean scores out of 100, same candidates measured both ways.
      GroupnWritten assessmentRecorded voice call
      Pasted at least once25269.074.6
      No paste detected42148.170.9
      Gap20.93.7

      On the written assessment, candidates who pasted outscored candidates who did not by 20.9 points. That is an enormous spread. It is the difference between the top of a shortlist and the reject pile.

      Put the same people on a recorded call and the gap falls to 3.7 points. 82.3 percent of the advantage disappears.

      The residual 3.7 point difference is statistically detectable, but it is smaller than the 9.8 point margin between our own two scoring passes. We would not build a hiring decision on it, and we are not asking you to.

      We checked whether this was an artefact of how we cleaned the data or which roles we included. It is not. Across five specifications, from the raw data through to the single largest role with every suspect score removed, the collapse lands between 82 and 85 percent.

      Sensitivity of the headline finding. All voice gaps significant at p below 0.001.
      SpecificationCandidatesWritten gapVoice gapGap closed
      Raw, all roles67820.83.782.1%
      Fallback scores removed67320.93.782.2%
      Fallback plus suspect cluster removed66721.43.982.0%
      Largest single role only, raw56121.03.284.7%
      Largest single role only, fully cleaned55621.43.384.6%

      How common is this? More common than most hiring managers assume

      Across all 699 written assessments:

      • 36.6 percent recorded at least one paste event.
      • 42.9 percent switched away from the assessment tab at least once.
      • 21.2 percent switched tabs five or more times. The highest single sitting recorded 54 switches.

      Tab switching behaves like a dose. A candidate leaving the tab once or twice scores about the same as one who never leaves. Past five switches, the score climbs sharply.

      Written score rises with tab switching0204060800 tab switchesn=39952.91-4 tab switchesn=15253.55-9 tab switchesn=6961.610+ tab switchesn=7968.7Mean written assessment score out of 100, by tab switches recorded during the sitting.
      Mean written assessment score by tab switches recorded during the sitting.
      Tab switchesnMean written score
      039952.9
      1-415253.5
      5-96961.6
      10+7968.7

      One detail complicates the lazy reading of this. Candidates who pasted spent longer on the assessment, not less: 28.9 minutes on average against 22.7 for candidates who did not. Whatever is happening, it is not people rushing through. It looks like effort, applied to the tool rather than to the answer.

      What the written test actually spreads people across

      The written scores look healthy in aggregate. The mean is 55.7, the median 60.0, and 36.2 percent of submissions clear 70. The distribution is wide, which is what you want from a screening instrument.

      Written assessment score distribution0%10%20%30%Under 4033.6%40 to 5412.2%55 to 6918.0%70 to 8425.0%85 or more11.2%Score out of 100. Share of 699 timed written submissions.
      Distribution of 699 timed written assessment scores.
      Score bandnShare
      Under 4023533.6%
      40 to 548512.2%
      55 to 6912618.0%
      70 to 8417525.0%
      85 or more7811.2%

      The problem is not that the spread is narrow. It is that a large part of it tracks browser behaviour rather than the candidate. A test that separates people cleanly is only useful if it separates them on something you meant to measure.

      What a recorded call measures that a written test cannot

      Voice answers are scored on five dimensions. Averaged across 1,848 recordings, they do not come out level.

      Where customer service candidates score strongest and weakestTone80.0%Accuracy73.6%Clarity73.2%Fluency70.4%Grammar68.6%Mean score as a share of the maximum for that dimension, across 1,848 recordings.
      Mean score on each rubric dimension.
      DimensionMean out of 20Share of maximum
      Tone16.080.0%
      Accuracy14.773.6%
      Clarity14.673.2%
      Fluency14.170.4%
      Grammar13.768.6%

      Tone is the strongest dimension at 80.0 percent, and grammar the weakest at 68.6 percent. The pattern is consistent: warmth, patience and the instinct to reassure an anxious customer show up reliably. Article use, tense agreement and preposition choice are where points go. In CEFR terms that is the signature of a strong B2 population, comfortable and effective in real conversation while still making visible grammatical slips.

      That ordering matters for anyone building a screening process, because it is the reverse of what a written test is good at. Written assessments measure grammar well and tone barely at all. They are strongest exactly where this candidate pool is weakest, and blind exactly where it is strongest. If you want to know more about which numbers actually predict customer service performance, our guide to contact center KPIs and metrics covers the operational side, and in-house versus outsourced contact centers covers where the screening burden lands.

      The constraint nobody screens for

      Voice work has a hardware floor that written work does not. We measure connection speed in the browser at assessment time.

      Applicant home bandwidthUnder 25 Mbpsn=23033.3%25 to 49n=27139.2%50 to 99n=17425.2%100 or moren=162.3%Measured in-browser at the time of the assessment. n=691.
      Measured home download speed, n=691.
      Download speednShare
      Under 25 Mbps23033.3%
      25 to 4927139.2%
      50 to 9917425.2%
      100 or more162.3%

      Median download speed is 36.4 Mbps and median latency 68 ms, which is comfortable for voice. But 33.3 percent measured under 25 Mbps, and only 2.3 percent reached 100 or more. Average typing speed was 43.3 words per minute, which matters for anyone expected to document a call while taking the next one.

      A candidate can be excellent and still be unable to hold a clean call from where they live. That is a solvable problem, but only if you measure it before you hire rather than after. Connectivity is one of the practical reasons teams cluster in particular cities, which we cover in our guide to call center outsourcing in the Philippines and on our Philippines contact center page.

      What we would take from this

      We are a staffing company, so treat the following as interested advice and check it against your own funnel.

      1. Assume an unproctored written score is partly a measure of tool access. Not entirely, and not for every candidate. But a 20.9 point spread that closes to 3.7 on a recorded call is not a small correction. The International Test Commission’s guidelines on computer-based and internet-delivered testing have warned about exactly this control gap for years, and they predate widely available generative tools.
      2. Record telemetry even if you do nothing with it. Paste events, tab switches and time on task cost nothing to capture and change how you read a score.
      3. Make candidates speak, early. A short recorded answer to a real scenario is the cheapest high-signal step available for a customer service role, and it is the step this data says survives contact with reality. This matters most for inbound voice work, where the first thirty seconds of a call set the outcome.
      4. Do not screen on self-declared English. We looked hard at whether self-described fluency predicted measured performance and could not separate the groups. Our scoring is not precise enough for us to publish that as a finding, which is itself the point: if we cannot measure the difference at this sample size, an application form dropdown certainly cannot.
      5. Test the connection before the offer, not after. A third of this pool would need an upgrade to run voice reliably.

      Reuse and citation

      All figures in this report may be reproduced with attribution and a link to this page. If you are citing it, this line works:

      Suggested citation
      Armasourcing, The Customer Service Screening Report 2026, based on 1,848 scored voice recordings and 699 written assessments from 670 customer service applicants.

      Underlying recordings and transcripts are not published. Every figure here is an aggregate, and no candidate is identifiable from anything in this report.

      Related first-party research: our Filipino Virtual Assistant Salary Report covers what candidates in this market ask to be paid. If you are working out how to structure a customer service team offshore, start with the contact center outsourcing hub, read how to choose a contact center outsourcing company, or compare building in-house against outsourcing.

      Frequently asked questions

      How many candidates does this report cover?

      670 customer service applicants, who between them produced 1,848 scored voice recordings and 699 timed written assessments between 2026-06-05 and 2026-07-27. The paired analysis behind the headline finding covers the 673 candidates who completed both.

      What share of candidates pasted into their assessment?

      36.6 percent of the 699 written assessments recorded at least one paste event, and 42.9 percent recorded at least one tab switch. Pasting is not proof of misconduct. Candidates paste prepared notes and rehearsed answers as well as anything else.

      Does a written assessment predict customer service performance?

      In this data it predicts less than it appears to. Candidates who pasted scored 20.9 points higher on the written test than those who did not, but only 3.7 points higher on a recorded voice call. 82.3 percent of the written advantage did not survive the change of format.

      Where are Filipino customer service candidates strongest and weakest?

      Across 1,848 scored recordings, tone is the strongest dimension at 80.0 percent of maximum and grammar the weakest at 68.6 percent. Warmth and customer handling instinct score consistently higher than grammatical precision.

      Is this data representative of the whole Philippine customer service market?

      No. 88 percent of the recordings come from a single hiring push for one United States moving and storage company, over roughly seven weeks. It is a deep look at one real funnel rather than a market-wide survey, and it should be cited that way.

      Other ways we can help

      See all services →
      Industry-trained

      Need a VA who already understands your industry?

      We don’t place generalists. Our VAs are matched and trained for the specific workflows of your sector.

      See industry VAs →
      Eli Gutilban - CEO of Armasourcing
      Written by

      Eli Gutilban

      CEO & Founder of Armasourcing

      Digital strategist with 10+ years of experience helping businesses scale with trained Filipino virtual assistants and managed contact center teams. Top Rated Plus on Upwork with 7,778+ verified hours and a 97% job success score.

      Book a Free Discovery Call

      Ready to Scale Your Business?

      Book a free discovery call and let us show you how we can help.

      Get My Free VA Match πŸ“… Book a Call Directly
      Matched Within a Week Top 3% Filipino Talent
      Call Hire Now