AI Text Detectors Fail Under Style Mimicry, Miss Rate Up to 48% in Scientific Writing
Decision Brief
Epoch AI tested three major AI text detectors: Pangram 3.3.2, GPTZero (model 2026-05-11-base), and Originality.ai Turbo 3.0.2. They built a corpus of 495 human-written texts (blogs, fiction, scientific writing) all predating ChatGPT. In simple-prompt detection, detectors were nearly perfect with false negative rates max 0.7%. However, when texts were generated by Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro mimicking specific author styles, 13% (38/297) of AI texts were missed on average, with Originality.ai missing 18%. Scientific writing was the weakest: Pangram missed 25%, GPTZero 24%, Originality.ai 29%. The worst-case combo was Gemini-generated academic paragraphs missed by Pangram at 48%; GPT-5.5 academic texts missed by Originality.ai at 39%. Despite different methods (neural network for Pangram, word predictability for GPTZero, statistical patterns for Originality.ai), trends were consistent. For educators and publishers, this means high accuracy on simple prompts is misleading—up to one-fifth of AI writing may evade detection in academic settings. Universities and journals relying on these tools must reassess their reliability, especially in scientific writing where detectors are most applied.
Sources
- The Decoder:AI News
- The Decoder:AI News
留言
登入后即可留言,和其他 builder 交换实测心得。
还没有留言,抢头香。