How well do AI Detectors work? And are general-purpose language models catching up? Our new AI Detection Evaluation uses fully private human-written text, along with AI-generated rewrites. We find that the age of specialized AI detectors like Pangram may be coming to an end.
16
11
268
49,605
GPT-6-Astra and Opus 5.5 exceed the performance of all specialized detectors at a fraction of the cost on accuracy (with an equal number of human and AI samples).

Oct 1, 2026 · 12:33 AM UTC

1
85
7,246
We find that all models are very strong at accurately labeling human text, with the top specialized and general models having perfect performance across our tests. The leaderboard below shows human-detection accuracy.
1
55
2,750
Weak overall accuracy is driven by mislabeling AI rewrites as human-written, a task for which models are less capable overall.
1
24
2,841
Additionally, we tested whether models were capable of modifying text to fool other AI-detectors and themselves. We found that Opus 5.5 and Astra can often rewrite more than 50% of a document without being detected as Mixed or AI by Pangram. These models can also fool themselves, though less successfully.
3
38
2,409
Human writing samples were sourced from private, pre-LLM documents sourced by the Vals team. AI writing samples were produced using a diverse pool of generator models. For our main evaluation, OpenAI and Anthropic latest-generation models are excluded from the generators for this data, because they’re tested as evaluators.
2
18
2,415
Sort replies: Relevant Recent Liked