{"id":229681,"date":"2026-08-14T08:49:00","date_gmt":"2026-08-14T08:49:00","guid":{"rendered":"https:\/\/www.9senses.ai\/?p=229681"},"modified":"2026-08-27T12:24:38","modified_gmt":"2026-08-27T12:24:38","slug":"the-broken-gate","status":"publish","type":"post","link":"https:\/\/www.9senses.ai\/de\/the-broken-gate\/","title":{"rendered":"AI detection: The broken gate"},"content":{"rendered":"<p>Teachers, editors and recruiters &#8211; and all of us who read online content &#8211; are struggling with the problem of telling AI-generated text from the &#8220;real thing,&#8221; something that was actually written by a human. Now, finally, detection seems to have caught up. In independent testing by Chicago Booth, AI detector Pangram was able to tell AI from human writing with an accuracy of more than 99.8%, and with almost no false positives. It was the only tool that remained reliable under the strict limits an academic institution would demand before accusing anyone of cheating with AI. And in early August 2026, Anthropic announced the arrival of a &#8220;made by AI&#8221; marker in content delivered by Claude models. All models launched on or after August 2, 2026, weave an invisible watermark into every text they generate &#8211; a mechanism Anthropic has since confirmed to be a version of Google DeepMind&#8217;s SynthID-Text &#8211; with older models to follow during the EU AI Act&#8217;s transition period. (Claude also signs supported file formats &#8211; for example images &#8211; with C2PA metadata; this article deals with text only.) There is no opting out, and, as the other big AI companies signed the same Code of Practice driven by the EU AI Act, the rest of the industry will likely follow soon. Google already did.<\/p>\n<p>It seems that the arms race is over. Teachers can grade again faithfully, editors can trust their inbox, and honest writers finally have proof on their side. Sounds great, doesn&#8217;t it?<\/p>\n<h3><span style=\"font-weight: normal\">The caveat: processing isn&#8217;t authoring<\/span><\/h3>\n<p>It isn&#8217;t great at all: take two students who hand in the same assignment; the first wrote every paragraph herself, then asked AI to improve the language, find typos and fix the grammar, while the second pasted the assignment straight into an AI chatbot, took the output and ran it through a &#8220;humanizer&#8221; tool, a service that exists abundantly online, and made some additional manual edits. In this new reality of watermarks and detectors, the first &#8211; honestly self-authored &#8211; essay carries an AI signature. The second comes back clean. The first one will likely have to defend herself; the second one goes unnoticed.<\/p>\n<p>Both AI detector and AI watermark measure only one thing: that a text was in contact with AI. The problem is that &#8220;contact with AI&#8221; is not fraud in itself. These two mechanisms tell honest from dishonest the same way a security scanner tells weapons from a few coins you forgot in your pocket. In the latter case, the error translates into a quick body search or another round trip through the detector; with written output, the result may be a misconduct hearing or &#8211; if it&#8217;s not a graded assessment but instead a job application or other writing &#8211; outright rejection without an opportunity to remedy the situation.<\/p>\n<p>And then there is the third person. Maybe a non-native speaker, maybe someone not so good with words, but someone with really original ideas worth sharing. Previously, this individual would have needed a lot of money to pay for a competent ghostwriter; nowadays, by feeding AI with their highly relevant thoughts and using the result, they can also be heard. And the question is: why shouldn&#8217;t they? Nothing they present is &#8220;AI slop,&#8221; as many people call AI-produced output without much value nowadays.<\/p>\n<h3><span style=\"font-weight: normal\">The fine print<\/span><\/h3>\n<aside class=\"ns-obj ns-obj--aside ns-obj--right ns-obj--c1 ns-obj--fmt-inverted\" style=\"width:63%\"><div class=\"ns-obj-body\"><h3>Watermarks in text are not safe<\/h3>\n<p>AI watermarks in a text that an LLM insert are removable. The standard format for many text documents, XML, doesn\u2019t know the concept of invisible and non-removable markers at all. The only thing Anthropic can technically do &#8211; and they state so clearly &#8211; is to weave certain patterns into the text or its formatting that can later be identified. For example, this could mean certain words used in a certain sequence that are not distorting the meaning but still provide proof of AI being at work.&nbsp;<\/p>\n<p>Changing the text sufficiently will weaken or remove those watermarks. Already now, the first \u201chumanizers\u201d are being launched that allegedly annihilate the exact protection Anthropic is using &#8211; a few days after the announcement! And more will follow.<\/p>\n<p>The bottom line: the person who really wants to disguise dishonest AI use will be able to do so, without much effort. The one who doesn\u2019t cheat and uses AI responsibly might not. This defeats the real purpose of detection. Who potentially gets caught is the innocent user, while the one with malicious intent knows all too well how to evade detection without much effort.<\/p><\/div><\/aside>\n<p>To Anthropic&#8217;s credit, the company says openly that a mark only means content may have been processed by Claude, not written by it. But will those who judge make that distinction? Will they evaluate whether something was just polished by AI or genuinely written by it? Have AI proofread your essay and it carries the mark, even though the model might not have added a single original thought, but just adjusted a few sentences, added a few commas and fixed the typos. But what&#8217;s worse: no mark proves nothing at all. The watermark confirms that a sequence of words was seen and touched by an AI model, and nothing about who did the thinking. The absence of a watermark, on the other hand, equally proves nothing: it could be created by an honest author with an original thought who hasn&#8217;t embraced AI yet &#8211; or the work of someone clever enough to let AI do the job and then disguise it.<\/p>\n<h3><span style=\"font-weight: normal\">The brittle Pangram success story<\/span><\/h3>\n<p>AI detectors are more of a problem than a solution and have always been. Most are too inaccurate to even be considered; and easily deceived. There is a new kid on the block that allegedly stands out: Pangram. We tested Pangram 3 previously and were not impressed, and we now repeated the evaluation with exactly this document using Pangram 4. The first six paragraphs were deliberately AI-generated by Claude Fable 5, but no content was truly &#8220;written&#8221; by AI; everything was based on a detailed prompt containing the writer&#8217;s ideas, the writer&#8217;s structure, the writer&#8217;s original thought delivered to Claude. &#8220;100% AI&#8221; was Pangram&#8217;s verdict. Five of those AI paragraphs were retained, one discarded. After 20 minutes of rewriting, where some words and a few sentences were altered, synonyms introduced, a few continuations and punctuation changed, Pangram confirmed: &#8220;100% human.&#8221; Nothing materially changed through all these passes. A statistical analysis shows that mostly words were altered, and we probably changed more than necessary because we simply liked ours better.&nbsp;<\/p>\n<figure class=\"ns-obj ns-obj--tablewrap ns-obj--full ns-obj--fmt-normal\">\n<table class=\"data\">\n<thead>\n<tr>\n<th>Pair<\/th>\n<th>Core statement<\/th>\n<th>Semantic<\/th>\n<th>Structural<\/th>\n<th>Word overlap &#8211; exact<\/th>\n<th>Word overlap &#8211; incl. synonyms<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>A1 -&gt; B1<\/td>\n<td>99%<\/td>\n<td>95%<\/td>\n<td>97%<\/td>\n<td>58.4%<\/td>\n<td>71.9%<\/td>\n<\/tr>\n<tr>\n<td>A2 -&gt; B2<\/td>\n<td>100%<\/td>\n<td>99%<\/td>\n<td>100%<\/td>\n<td>81.5%<\/td>\n<td>88.9%<\/td>\n<\/tr>\n<tr>\n<td>A3 -&gt; B3<\/td>\n<td>100%<\/td>\n<td>89%<\/td>\n<td>95%<\/td>\n<td>58.7%<\/td>\n<td>69.3%<\/td>\n<\/tr>\n<tr>\n<td>A4 -&gt; B4<\/td>\n<td>100%<\/td>\n<td>86%<\/td>\n<td>93%<\/td>\n<td>41.0%<\/td>\n<td>53.8%<\/td>\n<\/tr>\n<tr>\n<td>A5 -&gt; B6<\/td>\n<td>100%<\/td>\n<td>84%<\/td>\n<td>90%<\/td>\n<td>46.5%<\/td>\n<td>52.5%<\/td>\n<\/tr>\n<tr>\n<td>Mean<\/td>\n<td>99.8%<\/td>\n<td>90.6%<\/td>\n<td>95.0%<\/td>\n<td>57.2%<\/td>\n<td>67.3%<\/td>\n<\/tr>\n<\/tbody>\n<\/table><figcaption class=\"table-sources\">Core statement overlap: Similarity of the main claim or central argument, regardless of wording or supporting detail.<br \/>\nSemantic overlap: Similarity of the total substantive meaning, including secondary claims, examples and qualifications.<br \/>\nStructural overlap: Similarity in the order and organization of claims, examples, contrasts and conclusions.<br \/>\nExact word overlap: Literal overlap of normalized content words, excluding synonyms and broader paraphrases.<br \/>\nSynonym-adjusted word overlap: Exact word overlap plus clear word-family variants and conservative contextual synonyms.<\/figcaption><\/figure>\n<p>During this 20-minute round trip, first, Pangram labeled something that initially was entirely human-conceived and could never have been provided by an AI alone, as 100% AI, with our changes, it became human again. This is not a scientific analysis, and it will be hard to turn into one, because in the arms race between AI, humanizers and detectors, things are fluid.&nbsp;<\/p>\n<p>But our example shows that, while Pangram can detect pure AI-generated text reliably and defeat current humanizer tools well (but for how long?), it can be easily fooled, with very little human effort. But the key problem is another one: it cannot distinguish between AI generation and AI-based refinement, the former frowned upon and illegitimate, the latter perfectly defensible.<\/p>\n<p>Technically, it can&#8217;t be any other way. Detectors like Pangram can only do the reverse of the declaration Anthropic weaves into AI-created text: find patterns that are more likely to be AI-generated. With each new generation of AI models, this will become harder, as the language quality of those models is improving continuously, making detection inherently more and more brittle. And it will forever measure the wrong thing.<\/p>\n<h3><span style=\"font-weight: normal\">Detection is simply the wrong instrument<\/span><\/h3>\n<figure class=\"ns-obj ns-obj--related ns-obj--fmt-normal ns-ed-rail-item\"><div class=\"ns-vmap-editor-note\" style=\"position:relative;width:100%;overflow:hidden;aspect-ratio:5\/4;max-height:230px;\"><div style=\"position:absolute;inset:0;display:flex;align-items:center;justify-content:center;pointer-events:none\"><div style=\"padding:10px 12px;border:1px dashed rgba(170,180,205,.55);border-radius:6px;color:#8a93a6;font:600 12px\/1.4 sans-serif;background:rgba(8,13,25,.46)\"><div style=\"margin-bottom:6px\">Vector map<\/div><code style=\"display:inline-block;padding:3px 6px;border-radius:4px;background:rgba(127,140,170,.12);color:inherit\">[ninesenses_vectormap]<\/code><div style=\"margin-top:6px;font-weight:400\">Max height: 230px. Preview on the live page.<\/div><\/div><\/div><\/div><\/figure>\n<p>Let&#8217;s face it. Given these limitations, watermarking and detection are useless &#8211; and can even be counterproductive. Useless does not mean inaccurate: the technology itself can be impressively precise, and Pangram really does identify purely AI-generated text with high reliability. It is useless for the decision it is actually deployed for: judging honesty and authorship. Detection answers the question &#8220;has an AI model touched these words?&#8221; &#8211; but the teacher, editor or recruiter using it needs an answer to &#8220;who did the thinking, and was the work done legitimately?&#8221; Those are different questions, and no detection accuracy, however high, answers the latter one. It inconsistently measures intent and evasion effort rather than actual quality of what is delivered. What AI provides is interface quality (at least with the latest models), polished language, polished design. But what we actually must look at is content quality. And that has different properties depending on what we have in front of us. In real life, the quality of written prose is measured by originality and relevance. An editor wants a well-researched piece or a compelling fictional story; a recruiter wants to find a person capable of doing the job. If AI was used to polish how they present that, why should it matter? Before, they had someone proofread their manuscript too &#8211; or help from parents or friends to refine that job application. Content matters mostly, not format. Another issue we <a href=\"\/stealing\">covered elsewhere<\/a> is the copyright issue for creative work, and publishers might want proof of that too.<\/p>\n<aside class=\"ns-obj ns-obj--aside ns-obj--right ns-obj--c1 ns-obj--fmt-inverted\" style=\"width:69%\"><div class=\"ns-obj-body\"><h3>An uncomfortable truth for educators<\/h3>\n<p>Any institution that uses Pangram or any other checker for evaluating student honesty is misled, typically against the honest student. The watermark will make this worse, because it feels like proof &#8211; a confession signed by the machine itself, and irresistible to treat as one. It will be indicative of the wrong thing.<\/p>\n<p>The solution is simple yet difficult to implement. Teachers need to forget about detection. The form and quality of the final argument are nothing you can judge any longer, as long as you can&#8217;t rule out AI use. And this is only reliable if you were present when the output was created. This is a fact and nothing to debate, and as models get better and detectors and cheaters play cat and mouse, final output quality is no longer a signal, it is pure noise.<\/p>\n<p>The answer lies in what you judge. As soon as you ask students to document process steps, show and explain revisions from first draft to final product, you&#8217;re safe. In that context, allowing the right kind of documented AI use becomes helpful, because it teaches students what they will have to learn anyway. On top of that, you need visible demonstration of knowledge and capability, examined in a computer-free environment, where AI is absent. This means oral discourse, written exams, or, in other words: back to the old ways of evaluating capability.<\/p>\n<p>Forget detectors. Change the way you evaluate.<\/p><\/div><\/aside>\n<p>Where content matters least is in education. Writing an essay on Napoleon or George Washington was never about writing a novelty piece that has value to society. It was about <a href=\"\/the-lost-generation\">training the minds of students<\/a> on researching information, processing and structuring it, and for making an argument and presenting it coherently. The process matters, not the output. And that&#8217;s where education must go: disregard the level of polish for the output or create a level playing field by providing everyone with the ability for a final polishing pass using AI. But what becomes important is that students can demonstrate how they got to the result, with the ability to prove each step of the process. They also should be able to prove their ownership in oral and written examination. In such a world, the final form simply becomes irrelevant, and so do AI detectors.<\/p>\n<p><\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI detectors and the new Claude watermark measure contact with AI, and are easy to evade. Honest AI users get marked; those who really cheat walk easily. We discuss why detection is the wrong approach, and what could actually work.<\/p>","protected":false},"author":15,"featured_media":229755,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"ns_references":"<h3>Research and Data<\/h3>\r\n<p><em>Dathathri, S., et al. (2024). <a href=\"https:\/\/doi.org\/10.1038\/s41586-024-08025-4\" target=\"_blank\" rel=\"noopener\">Scalable watermarking for identifying large language model outputs<\/a>. Nature, 634, 818\u2013823.<\/em><\/p>\r\n<p><em>Han, X., Li, Q., Ni, J., &amp; Zulkernine, M. (2025). <a href=\"https:\/\/arxiv.org\/abs\/2508.20228\" target=\"_blank\" rel=\"noopener\">Robustness assessment and enhancement of text watermarking for Google's SynthID<\/a>. arXiv preprint arXiv:2508.20228.<\/em><\/p>\r\n<p><em>Hwang, D., Woo, S., Gao, T., Luo, R., &amp; Baek, S. (2024). <a href=\"https:\/\/arxiv.org\/abs\/2412.12511\" target=\"_blank\" rel=\"noopener\">Invisible watermarks: Attacks and robustness<\/a>. arXiv preprint arXiv:2412.12511.<\/em><\/p>\r\n<p><em>Jabarian, B., &amp; Imas, A. (2025). <a href=\"https:\/\/bfi.uchicago.edu\/working-papers\/artificial-writing-and-automated-detection\/\" target=\"_blank\" rel=\"noopener\">Artificial writing and automated detection<\/a>. Becker Friedman Institute Working Paper No. 2025-116. DOI: 10.2139\/ssrn.5407424.<\/em><\/p>\r\n<p><em>Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., &amp; Zou, J. (2023). <a href=\"https:\/\/doi.org\/10.1016\/j.patter.2023.100779\" target=\"_blank\" rel=\"noopener\">GPT detectors are biased against non-native English writers<\/a>. Patterns, 4(7), Article 100779. DOI: 10.1016\/j.patter.2023.100779.<\/em><\/p>\r\n<p><em>Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., &amp; Khuat, H. Q. (2024). <a href=\"https:\/\/arxiv.org\/abs\/2403.19148\" target=\"_blank\" rel=\"noopener\">GenAI detection tools, adversarial techniques and implications for inclusivity in higher education<\/a>. arXiv preprint arXiv:2403.19148.<\/em><\/p>\r\n<p><em>Van Vlasselaer, M., Van Droogenbroeck, F., &amp; Spruyt, B. (2026). <a href=\"https:\/\/doi.org\/10.1007\/s40979-026-00226-w\" target=\"_blank\" rel=\"noopener\">Who wrote this? Evaluating the reliability of AI detection tools in higher education<\/a>. International Journal for Educational Integrity, 22, Article 16. DOI: 10.1007\/s40979-026-00226-w.<\/em><\/p>\r\n\r\n<h3>Policy and Provenance<\/h3>\r\n<p><em>Anthropic. (2026, August 11). <a href=\"https:\/\/support.claude.com\/en\/articles\/16266773-how-claude-marks-ai-generated-content\" target=\"_blank\" rel=\"noopener\">How Claude marks AI-generated content<\/a>. Anthropic Help Center.<\/em><\/p>\r\n<p><em>Anthropic. (2026, August 14). <a href=\"https:\/\/www.anthropic.com\/news\/claude-text-watermark\" target=\"_blank\" rel=\"noopener\">How Claude\u2019s text watermarking works<\/a>.<\/em><\/p>\r\n<p><em>Bridgeman, A., &amp; Liu, D. (2024). <a href=\"https:\/\/educational-innovation.sydney.edu.au\/teaching@sydney\/frequently-asked-questions-about-the-two-lane-approach-to-assessment-in-the-age-of-ai\/\" target=\"_blank\" rel=\"noopener\">Frequently asked questions about the two-lane approach to assessment in the age of generative AI<\/a>. The University of Sydney.<\/em><\/p>\r\n<p><em>Coalition for Content Provenance and Authenticity. <a href=\"https:\/\/c2pa.org\/specifications\/\" target=\"_blank\" rel=\"noopener\">Content Credentials: C2PA technical specification<\/a>.<\/em><\/p>\r\n<p><em>European Commission. (2026). <a href=\"https:\/\/digital-strategy.ec.europa.eu\/en\/policies\/guidelines-transparency-ai-generated-content\" target=\"_blank\" rel=\"noopener\">Guidelines on transparency obligations for providers and deployers of certain AI systems<\/a>. Artificial Intelligence Act, Article 50.<\/em><\/p>\r\n<p><em>European Commission. (2026, June 10). <a href=\"https:\/\/digital-strategy.ec.europa.eu\/en\/news\/commission-publishes-code-practice-marking-and-labelling-ai-generated-content\" target=\"_blank\" rel=\"noopener\">Code of Practice on Transparency of AI-Generated Content<\/a>.<\/em><\/p>\r\n<p><em>Google DeepMind. (2024, May 14). <a href=\"https:\/\/deepmind.google\/blog\/watermarking-ai-generated-text-and-video-with-synthid\/\" target=\"_blank\" rel=\"noopener\">Watermarking AI-generated text and video with SynthID<\/a>.<\/em><\/p>\r\n<p><em>Turnitin. (2023). <a href=\"https:\/\/www.turnitin.com\/blog\/understanding-false-positives-within-our-ai-writing-detection-capabilities\" target=\"_blank\" rel=\"noopener\">Understanding false positives within our AI writing detection capabilities<\/a>.<\/em><\/p>\r\n<p><em>Vanderbilt University. (2023, August 16). <a href=\"https:\/\/www.vanderbilt.edu\/brightspace\/2023\/08\/16\/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector\/\" target=\"_blank\" rel=\"noopener\">Guidance on AI detection and why Turnitin\u2019s AI detector was disabled<\/a>.<\/em><\/p>\r\n\r\n<h3>Commentary and Industry Reporting<\/h3>\r\n<p><em>Chicago Booth Review. (2025). <a href=\"https:\/\/www.chicagobooth.edu\/review\/2025\/december\/do-ai-detectors-work-well-enough-trust\" target=\"_blank\" rel=\"noopener\">Do AI detectors work well enough to trust?<\/a><\/em><\/p>\r\n<p><em>Pangram Labs. <a href=\"https:\/\/www.pangramlabs.com\/\" target=\"_blank\" rel=\"noopener\">All about false positives in AI detectors; Third-party Pangram evaluations<\/a>. Company blog.<\/em><\/p>\r\n<p><em>Requarth, T. (2026). <a href=\"https:\/\/timrequarth.substack.com\/\" target=\"_blank\" rel=\"noopener\">Why you shouldn\u2019t trust AI detectors<\/a>. Substack newsletter.<\/em><\/p>","ns_references_title":"References and further reading","n9tr_seo_title_de_DE":"KI-Erkennung: Der kaputte Filter - 9senses.ai","n9tr_seo_description_de_DE":"KI-Detektoren und das neue Claude-Wasserzeichen messen Kontakt mit KI und sind leicht zu umgehen. Ehrliche Nutzer werden markiert, echte Betr\u00fcger kommen durch. Warum Erkennung der falsche Ansatz ist.","n9tr_seo_title_fr_FR":"","n9tr_seo_description_fr_FR":"","footnotes":""},"categories":[47],"tags":[],"class_list":["post-229681","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog"],"acf":{"tag_line":"Why AI detection flags the honest and misses the fraudsters","about":"AI detectors are better than ever, and Claude now watermarks AI output at the source. But both measure contact with AI, not fraud: honest writers who use AI to polish their own work get flagged, while cheaters launder their text clean with limited effort. We look at the evidence, the arithmetic of false accusations, and what would actually restore trust.","tldr":"A detector verdict or watermark can only ever mean \"AI touched this text\", never \"AI wrote it\" - it is nothing but context and never proof. Building important decisions on these foundations is risky and can be thoroughly unfair. Instead, we need different approaches for different situations that evaluate what really matters.","ai_support":"The original ideas, structure, and much of the language are human-created, but AI was used to develop, enrich, or rework portions of the content \u2014 for example, researching sources, rewriting sections for clarity, or expanding on arguments."},"_links":{"self":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/posts\/229681","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/users\/15"}],"replies":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/comments?post=229681"}],"version-history":[{"count":42,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/posts\/229681\/revisions"}],"predecessor-version":[{"id":230555,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/posts\/229681\/revisions\/230555"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/media\/229755"}],"wp:attachment":[{"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/media?parent=229681"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/categories?post=229681"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.9senses.ai\/de\/wp-json\/wp\/v2\/tags?post=229681"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}