Diesen Beitrag anhören

0:00
Why AI detection flags the honest and misses the fraudsters

KI-Erkennung: Der kaputte Filter

AI detectors and the new Claude watermark measure contact with AI, and are easy to evade. Honest AI users get marked; those who really cheat walk easily. We discuss why detection is the wrong approach, and what could actually work.

Teachers, editors and recruiters – and all of us who read online content – are struggling with the problem of telling AI-generated text from the “real thing,” something that was actually written by a human. Now, finally, detection seems to have caught up. In independent testing by Chicago Booth, AI detector Pangram was able to tell AI from human writing with an accuracy of more than 99.8%, and with almost no false positives. It was the only tool that remained reliable under the strict limits an academic institution would demand before accusing anyone of cheating with AI. And in early August 2026, Anthropic announced the arrival of a “made by AI” marker in content delivered by Claude models. All models launched on or after August 2, 2026, weave an invisible watermark into every text they generate – a mechanism Anthropic has since confirmed to be a version of Google DeepMind’s SynthID-Text – with older models to follow during the EU AI Act’s transition period. (Claude also signs supported file formats – for example images – with C2PA metadata; this article deals with text only.) There is no opting out, and, as the other big AI companies signed the same Code of Practice driven by the EU AI Act, the rest of the industry will likely follow soon. Google already did.

It seems that the arms race is over. Teachers can grade again faithfully, editors can trust their inbox, and honest writers finally have proof on their side. Sounds great, doesn’t it?

The caveat: processing isn’t authoring

It isn’t great at all: take two students who hand in the same assignment; the first wrote every paragraph herself, then asked AI to improve the language, find typos and fix the grammar, while the second pasted the assignment straight into an AI chatbot, took the output and ran it through a “humanizer” tool, a service that exists abundantly online, and made some additional manual edits. In this new reality of watermarks and detectors, the first – honestly self-authored – essay carries an AI signature. The second comes back clean. The first one will likely have to defend herself; the second one goes unnoticed.

Both AI detector and AI watermark measure only one thing: that a text was in contact with AI. The problem is that “contact with AI” is not fraud in itself. These two mechanisms tell honest from dishonest the same way a security scanner tells weapons from a few coins you forgot in your pocket. In the latter case, the error translates into a quick body search or another round trip through the detector; with written output, the result may be a misconduct hearing or – if it’s not a graded assessment but instead a job application or other writing – outright rejection without an opportunity to remedy the situation.

And then there is the third person. Maybe a non-native speaker, maybe someone not so good with words, but someone with really original ideas worth sharing. Previously, this individual would have needed a lot of money to pay for a competent ghostwriter; nowadays, by feeding AI with their highly relevant thoughts and using the result, they can also be heard. And the question is: why shouldn’t they? Nothing they present is “AI slop,” as many people call AI-produced output without much value nowadays.

Das Kleingedruckte

To Anthropic’s credit, the company says openly that a mark only means content may have been processed by Claude, not written by it. But will those who judge make that distinction? Will they evaluate whether something was just polished by AI or genuinely written by it? Have AI proofread your essay and it carries the mark, even though the model might not have added a single original thought, but just adjusted a few sentences, added a few commas and fixed the typos. But what’s worse: no mark proves nothing at all. The watermark confirms that a sequence of words was seen and touched by an AI model, and nothing about who did the thinking. The absence of a watermark, on the other hand, equally proves nothing: it could be created by an honest author with an original thought who hasn’t embraced AI yet – or the work of someone clever enough to let AI do the job and then disguise it.

The brittle Pangram success story

AI detectors are more of a problem than a solution and have always been. Most are too inaccurate to even be considered; and easily deceived. There is a new kid on the block that allegedly stands out: Pangram. We tested Pangram 3 previously and were not impressed, and we now repeated the evaluation with exactly this document using Pangram 4. The first six paragraphs were deliberately AI-generated by Claude Fable 5, but no content was truly “written” by AI; everything was based on a detailed prompt containing the writer’s ideas, the writer’s structure, the writer’s original thought delivered to Claude. “100% AI” was Pangram’s verdict. Five of those AI paragraphs were retained, one discarded. After 20 minutes of rewriting, where some words and a few sentences were altered, synonyms introduced, a few continuations and punctuation changed, Pangram confirmed: “100% human.” Nothing materially changed through all these passes. A statistical analysis shows that mostly words were altered, and we probably changed more than necessary because we simply liked ours better. 

Pair Core statement Semantic Structural Word overlap – exact Word overlap – incl. synonyms
A1 -> B1 99% 95% 97% 58.4% 71.9%
A2 -> B2 100% 99% 100% 81.5% 88.9%
A3 -> B3 100% 89% 95% 58.7% 69.3%
A4 -> B4 100% 86% 93% 41.0% 53.8%
A5 -> B6 100% 84% 90% 46.5% 52.5%
Mean 99.8% 90.6% 95.0% 57.2% 67.3%
Core statement overlap: Similarity of the main claim or central argument, regardless of wording or supporting detail.
Semantic overlap: Similarity of the total substantive meaning, including secondary claims, examples and qualifications.
Structural overlap: Similarity in the order and organization of claims, examples, contrasts and conclusions.
Exact word overlap: Literal overlap of normalized content words, excluding synonyms and broader paraphrases.
Synonym-adjusted word overlap: Exact word overlap plus clear word-family variants and conservative contextual synonyms.

During this 20-minute round trip, first, Pangram labeled something that initially was entirely human-conceived and could never have been provided by an AI alone, as 100% AI, with our changes, it became human again. This is not a scientific analysis, and it will be hard to turn into one, because in the arms race between AI, humanizers and detectors, things are fluid. 

But our example shows that, while Pangram can detect pure AI-generated text reliably and defeat current humanizer tools well (but for how long?), it can be easily fooled, with very little human effort. But the key problem is another one: it cannot distinguish between AI generation and AI-based refinement, the former frowned upon and illegitimate, the latter perfectly defensible.

Technically, it can’t be any other way. Detectors like Pangram can only do the reverse of the declaration Anthropic weaves into AI-created text: find patterns that are more likely to be AI-generated. With each new generation of AI models, this will become harder, as the language quality of those models is improving continuously, making detection inherently more and more brittle. And it will forever measure the wrong thing.

Detection is simply the wrong instrument

Let’s face it. Given these limitations, watermarking and detection are useless – and can even be counterproductive. Useless does not mean inaccurate: the technology itself can be impressively precise, and Pangram really does identify purely AI-generated text with high reliability. It is useless for the decision it is actually deployed for: judging honesty and authorship. Detection answers the question “has an AI model touched these words?” – but the teacher, editor or recruiter using it needs an answer to “who did the thinking, and was the work done legitimately?” Those are different questions, and no detection accuracy, however high, answers the latter one. It inconsistently measures intent and evasion effort rather than actual quality of what is delivered. What AI provides is interface quality (at least with the latest models), polished language, polished design. But what we actually must look at is content quality. And that has different properties depending on what we have in front of us. In real life, the quality of written prose is measured by originality and relevance. An editor wants a well-researched piece or a compelling fictional story; a recruiter wants to find a person capable of doing the job. If AI was used to polish how they present that, why should it matter? Before, they had someone proofread their manuscript too – or help from parents or friends to refine that job application. Content matters mostly, not format. Another issue we covered elsewhere is the copyright issue for creative work, and publishers might want proof of that too.

Where content matters least is in education. Writing an essay on Napoleon or George Washington was never about writing a novelty piece that has value to society. It was about training the minds of students on researching information, processing and structuring it, and for making an argument and presenting it coherently. The process matters, not the output. And that’s where education must go: disregard the level of polish for the output or create a level playing field by providing everyone with the ability for a final polishing pass using AI. But what becomes important is that students can demonstrate how they got to the result, with the ability to prove each step of the process. They also should be able to prove their ownership in oral and written examination. In such a world, the final form simply becomes irrelevant, and so do AI detectors.

Referenzen und weitere Quellen

Forschung und Daten

Dathathri, S., et al. (2024). Scalable watermarking for identifying large language model outputs. Nature, 634, 818–823.

Han, X., Li, Q., Ni, J., & Zulkernine, M. (2025). Robustness assessment and enhancement of text watermarking for Google's SynthID. arXiv preprint arXiv:2508.20228.

Hwang, D., Woo, S., Gao, T., Luo, R., & Baek, S. (2024). Invisible watermarks: Attacks and robustness. arXiv preprint arXiv:2412.12511.

Jabarian, B., & Imas, A. (2025). Artificial writing and automated detection. Becker Friedman Institute Working Paper No. 2025-116. DOI: 10.2139/ssrn.5407424.

Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. DOI: 10.1016/j.patter.2023.100779.

Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). GenAI detection tools, adversarial techniques and implications for inclusivity in higher education. arXiv preprint arXiv:2403.19148.

Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. International Journal for Educational Integrity, 22, Article 16. DOI: 10.1007/s40979-026-00226-w.

Policy and Provenance

Anthropic. (2026, August 11). How Claude marks AI-generated content. Anthropic Help Center.

Anthropic. (2026, August 14). How Claude’s text watermarking works.

Bridgeman, A., & Liu, D. (2024). Frequently asked questions about the two-lane approach to assessment in the age of generative AI. The University of Sydney.

Coalition for Content Provenance and Authenticity. Content Credentials: C2PA technical specification.

European Commission. (2026). Guidelines on transparency obligations for providers and deployers of certain AI systems. Artificial Intelligence Act, Article 50.

European Commission. (2026, June 10). Code of Practice on Transparency of AI-Generated Content.

Google DeepMind. (2024, May 14). Watermarking AI-generated text and video with SynthID.

Turnitin. (2023). Understanding false positives within our AI writing detection capabilities.

Vanderbilt University. (2023, August 16). Guidance on AI detection and why Turnitin’s AI detector was disabled.

Kommentare und Berichte zu Unternehmen

Chicago Booth Review. (2025). Do AI detectors work well enough to trust?

Pangram Labs. All about false positives in AI detectors; Third-party Pangram evaluations. Company blog.

Requarth, T. (2026). Why you shouldn’t trust AI detectors. Substack newsletter.