21 September 2026
AI detectors cannot tell who wrote a CV. Screen this instead
Seven detectors flagged 61% of human-written TOEFL essays as AI. A 2026 study of 13 detectors found false positive rates from 0% to 100%. Here is what to screen instead of authorship.
Every hiring team has now had the conversation. Applications are up, half of them read like the same confident paragraph, and somebody asks whether there is a tool that spots the ones written by ChatGPT. There is. Several, in fact. The evidence says they cannot do the job, and that wiring one into a rejection rule creates a bigger problem than the one it solves.
This is not an argument for screening less. It is an argument for screening the right object: what a candidate claims and whether the document evidences it, rather than who or what typed the sentences.
What people mean by "an AI-written CV"
Three different things get filed under one complaint. A candidate who used a model to tidy their phrasing. A candidate who generated the whole document from a job advert, including experience they do not have. And a mass applicant firing hundreds of tailored variants at every requisition in a sector.
Only the second and third are hiring problems, and neither is an authorship problem. The first is spellcheck with better manners. A detector cannot separate them, because all three produce the same surface statistics.
The evidence that started this: 61% false positives
The foundational study is GPT detectors are biased against non-native English writers. Running seven widely used detectors over TOEFL essays written by humans, the authors found the detectors misclassified "over half of the TOEFL essays as 'AI-generated' (average false positive rate: 61.22%)".
The individual numbers are worse than the average suggests. "All seven detectors unanimously identified 18 of the 91 TOEFL essays (19.78%) as AI-generated", and "89 of the 91 TOEFL essays (97.80%) are flagged as AI-generated by at least one detector". Against US eighth-grade essays written by native speakers, the same detectors showed near-perfect accuracy.
In other words: the tool works on the population that does not need protecting and fails on the one that does. In a hiring funnel, that is not a quality problem. It is a protected-characteristic problem wearing a quality problem's clothes.
Why the bias is structural
The mechanism is text perplexity — roughly, how surprising each next word is. Detectors treat low perplexity as evidence of machine authorship. Writers with a smaller active vocabulary in English produce lower-perplexity prose, and the study found the association highly significant (P = 9.74E-05).
The clinching experiment: when the authors prompted a model to rewrite the same TOEFL essays to "sound more like that of a native speaker", the average false positive rate fell from 61.22% to 11.77%. Making the text more machine-processed made it look more human to the detectors.
The effect is not confined to essays. Across 1,574 ICLR 2023 paper abstracts, authors from non-native-English countries wrote significantly lower-perplexity abstracts (P = 0.035). Published academics, flagged by the same logic.
The 2026 update: the spread is 0% to 100%
A larger study, Style as a Confound, submitted on 27 August 2026, tested 13 detectors across 135,389 document pairs spanning 2018 to 2025. False positive rates on human-written text ranged "from 0.0% to 100.0%" depending on the detector and the writing style.
The paper also records that "the same edits increased AI scores in some detectors but decreased them in others". Two vendors, one document, opposite verdicts. If your process rejects on a score, the outcome for a real applicant depends on which contract your procurement team happened to sign.
Stacking detectors makes it worse, not better
The instinct when one tool looks unreliable is to require agreement from two. The numbers rule that out. In the TOEFL study, 97.80% of the human-written essays were flagged by at least one of the seven detectors, while only 19.78% were flagged by all seven. Loosen the rule to "flagged by any tool" and you reject almost everyone; tighten it to unanimity and you still reject one in five people who wrote every word themselves.
There is no threshold in between that makes the underlying signal valid, because the signal is measuring fluency, not authorship.
The volume pressure that makes detection tempting
The instinct is understandable, because the inbox is genuinely worse. According to Ashby's analysis published on 7 May 2026, drawing on over 100 million applications and 200,000 jobs, “applications per hire have tripled since 2021, with roles now receiving more than 300 applications per hire on average”, and “candidates today are roughly 50% less likely to receive an interview than they were five years ago”.
Volume is the real problem. Detection is a proxy that punishes the wrong people for it.
What an auto-reject rule actually exposes you to
Rejecting on a detector score is an automated decision made on a signal the research says correlates with the writer's first language. Three live obligations meet it.
New York City. Local Law 144 requires a bias audit of an automated employment decision tool "within one year of the use of the tool", publication of the audit results, and notice to candidates at least 10 business days before use. DCWP has been enforcing since 5 July 2023. A detector used to filter applicants is squarely the kind of tool the law describes.
The EU. Transparency obligations under Article 50 of the AI Act apply from 2 August 2026. The high-risk obligations covering employment under Annex III were pushed back to 2 December 2027 by the digital omnibus package, per the AI Act Explorer's tracker. The later date is a delay on the heavier duties, not a licence in the meantime.
Candidates. Greenhouse's survey of 2,950 job seekers, published on 1 May 2026, found 70% were never told upfront that AI would evaluate them and 57% think disclosure should be legally required. A rejection you cannot explain is a complaint you cannot answer.
The related trap — believing you have decision support when a regulator would see solely automated decision-making — is covered in our piece on the AI CV screening compliance line.
The attack nobody is screening for
While detectors chase authorship, a real adversarial problem has arrived. Researchers analysing roughly 200,000 real resumes found that about 1% contained hidden prompt injections aimed at the screening model, according to a study posted on 27 May 2026. More than 90% used no explicit instruction — no "ignore previous instructions", just text engineered to tilt the model's judgement.
One in a hundred is a lot when a requisition draws three hundred applications. And notice the design implication: a screener whose scores are produced by a language model is manipulable by the document it is reading. A screener where the model only extracts quoted facts, and deterministic code does the scoring, is not — the worst an injected line can do is misquote itself into a human's view.
Screen claims, not prose
The replacement for authorship detection is evidence extraction. For each requirement in the advert, ask one question: does this document contain a specific, checkable claim that meets it, and where exactly?
- Quote the line. Every score should carry the sentence it came from, so a human can verify it in one click and a rejected candidate can be given a reason.
- Score deterministically. Let the model extract facts and quotes; let code do the arithmetic. The same CV against the same advert then scores the same every time, which is what an audit needs.
- Flag, never auto-reject. A missing requirement is a label a person can overrule, not a deletion.
- Test the claims later, not the prose. Polished writing is verified at interview, in a work sample, or in a reference. That is where it was always verified.
- Ignore style entirely. Spelling, grammar and the school someone attended are proxies for class and background, not for competence.
A 30-minute audit of your current stack
Ask your ATS vendor four questions and write the answers down. Does any part of the pipeline score, rank or filter on a machine-authorship or writing-quality signal? If a candidate asks why they were rejected, what artefact do we show them? Where does the score come from — a model's judgement, or code operating on extracted facts? And has any automated tool in the funnel had a bias audit, with results we can publish?
If the answers are uncomfortable, that is the finding. Detection promises to remove work from a human and quietly moves legal risk onto them instead.
HireSieve is built to the opposite rule: a model extracts facts and quotes, code does every calculation, every score quotes the CV line it came from, and nothing auto-rejects. It will not guess whether a CV was written by AI, because — on the published evidence — nothing reliably can.