AI detector percentages can look precise, but recent research shows that they can capture features of writing that do not map neatly to authorship. An arXiv preprint submitted in August 2026 tested published English abstracts from four domains and compared older work from 2013 to 2015 with newer work from 2023 to 2025. The results show why a detector score should be read cautiously, especially when a writer has revised or refined text rather than generated a passage from scratch.
The arXiv study used light refine-the-abstract-only edits as a proxy for guideline-compliant AI assistance. Those lightly edited abstracts were flagged at rates between 38 and 80 per cent. By comparison, recent unmodified abstracts from 2023 to 2025 were flagged at rates between 9 and 15 per cent. The researchers also found significantly higher false positive rates in non-STEM domains than in STEM domains, with p less than 0.001. This means even limited editing could sharply change a detector result.
Pangram reported very different figures from third-party evaluations. A Vrije Universiteit Brussel evaluation in June 2026 reported 97.5 per cent detection on fully AI-generated text, 95 per cent on humanised text and zero per cent false positives on essays by people using English as a second language. Other evaluations reported similarly high detection. These numbers do not contradict the arXiv study directly because the evaluations and the academic study measured different text sets, conditions and outcomes.
That gap matters for anyone who uses a rewriting or paraphrasing tool on text they originally wrote. The arXiv study found that lightly refined abstracts could receive much higher detector scores than recent unmodified abstracts. It also found that elevated scores tracked long-token density and Academic Word List density, not authorship intent alone. A writer can therefore make ordinary changes to wording and still see a result that looks more suspicious, even when the detector cannot establish why those changes were made.
A percentage from an AI detector is not evidence of intent. It does not show why a passage has particular linguistic features, whether a writer generated text, or whether a writer used limited assistance while revising original work. In the arXiv study, detection after humanisation fell below 4 per cent and false negative rates rose above 96 per cent. The authors conclude that detector scores alone should not determine misconduct cases. That conclusion puts the focus on evidence and process rather than one numerical result.
A practical habit for honest writers is to keep records of how a piece of work developed. Save early drafts, retain version history where it is available and keep notes that show how ideas, sources and wording changed over time. These records do not change a detector score, but they can help demonstrate a real writing process if questions arise later. For students, researchers and professional writers, that kind of evidence can provide useful context that a detector percentage cannot supply on its own.
The safest approach is to follow the disclosure policy set by the relevant school, university or employer. If a policy allows certain forms of rewriting or AI assistance, use them only within those rules and keep evidence of your own work. If disclosure is required, make it in the form the institution expects. Detector research shows that percentages can behave differently across tasks and text types, so they should not replace clear policies, documented writing processes and direct evidence about how a piece of work was produced.
Tags
Ready to put this into practice?
Use our free AI writing tools to apply what you just learned — join 2M+ students today.
Try Free Tools Now