Digital Economy

Texts written by AI, because the new controls can be circumvented

OpenAI is also starting to attribute the texts produced by its artificial intelligence to distinguish them from those written by humans.

FILE PHOTO: A keyboard is placed in front of a displayed OpenAI logo in this illustration taken February 21, 2023. REUTERS/Dado Ruvic/Illustration/File Photo REUTERS

5' min read

Translated by AI
Versione italiana

5' min read

Translated by AI
Versione italiana

OpenAI is also beginning to watermark the texts generated by its artificial intelligence to distinguish them from those written by humans. Following Anthropic, which announced a watermark for Claude in August, OpenAI announced on 5 October that, in the coming weeks, it will introduce an invisible mark in texts produced by ChatGPT and Codex within the European Union. However, the company’s own published findings show just how fragile this traceability is: changing just a few words can be enough to allow much of the content examined to slip through the net.

For businesses, schools and publishers, the prospect is an attractive one: to have an indication of a document’s origin, without relying solely on the statements of the person submitting it. The risk, however, is attributing a value to that indication that it does not possess. One text may have been produced entirely by a machine and passed the check; another may have been created by a person but retain traces of subsequent automated reworking.

Loading...

The impetus comes from the AI Act. The European transparency provisions, which came into force on 2 August, require providers to ensure that synthetic content is recognisable by automated tools, with certain exceptions, including specific functions designed solely to assist with editing. For systems already on the market prior to that date, the transitional deadline for labelling is 2 December 2026. The obligation creates a shared incentive to develop these tools, but does not eliminate their technical weaknesses.

A text watermark is not a hidden label within the document. It does not consist of invisible characters that can be deleted, nor of information that disappears when a passage is copied into another programme. As illustrated by the study on SynthID-Text published in *Nature* in 2024, the watermarking takes place during generation: from among the various plausible continuations of a sentence, the system makes its choice according to a mechanism that leaves a recognisable statistical pattern.

The reader sees a normal sequence of words. The detector, on the other hand, looks for that pattern across the entire text. This is a crucial distinction: changing the font or pasting the text without formatting does not alter the words and therefore does not, in itself, remove the signature.

Anthropic uses a variant of SynthID-Text and has opted for global deployment, explaining that it does not yet have a reliable way of restricting its use to specific geographical areas. The signature does not identify the user. Furthermore, the check looks for the ‘Claude’ signal: it is not a universal verifier capable of recognising any artificial intelligence.

This is where the problems arise. The first way to circumvent the checks is to manipulate the words themselves. In the evaluations published by OpenAI, based on English responses taken from questions in the Eli5 dataset, the detection rate for 400-token segments — the units of text processed by the model — drops from around 92 per cent to 66 per cent when 10 per cent of the words are replaced with synonyms. With 25 per cent of words replaced, it plummets to 17 per cent.

A second method is paraphrasing: retaining the content whilst changing the wording, sentence structure and order of presentation. This can be carried out by another AI. The text thus remains of automated origin, whilst the trace of the first generation becomes less evident.

The issue is documented in the study by Vinu Sankar Sadasivan and colleagues, ‘Can AI-Generated Text be Reliably Detected?’, published in *Transactions on Machine Learning Research* and updated in 2025. The authors experimented with successive paraphrasing of passages of around 300 tokens, finding that the detection accuracy of various systems—including those based on watermarks—declined, often with only a slight deterioration in quality.

Translation introduces another layer of transformation: words and sentence structures change. However, a service that rewrites or translates text may imprint a new AI signature. Anthropic points out, for example, that a translation produced by Claude is flagged as such. It is therefore possible to lose track of the original model and pick up the second one’s signature instead.

Loading...

Complicating matters is the confusion between watermarks and commercial detectors. The latter can search for recurring features in automated writing without having access to the watermark key. They assess how closely a text resembles those produced by the models: their verdict therefore also depends on the language, the type of document and the author’s style.

The study by Weixin Liang and colleagues, published in *Patterns* in 2023, highlighted both sides of the issue. The detectors examined tended to misclassify English texts produced by non-native speakers as artificial. At the same time, stylistic reworking of content generated by ChatGPT reduced its recognition.

There are, however, more recent tools, such as Pangram, which aim to detect machine-generated text even without watermarks, using models trained to distinguish between human, automated and mixed text. In tests reported by the company, Pangram 4 identifies AI intervention in 98.83 per cent of texts reworked by 13 services that claim to make them ‘human’. Pangram also indicates in detail which sentences in a text are human-written and which are AI-generated.

The company states, however, that it has been calibrated to favour false negatives over false positives. In other words, rather than labelling human-generated content as artificial – aware of the stigma that such a judgement may entail – it prefers to err on the other side.

The issue is complicated by the fact that the line between AI-generated writing and human writing is a fine one. There are texts written by one AI and then edited by another. Others are written by a human, but the style is influenced by that of the AI (as is already the case in many instances).

Finally, there remains the problem of evaluating these systems from the outside. OpenAI is initially restricting access to its detector to approved researchers and specialist organisations.

In the September 2026 preprint *Watermarks Without Verification*, Alexander Nemecek and colleagues investigated an open-source implementation of SynthID-Text on two publicly available weight-based models, precisely because they were unable to directly verify commercial systems. The authors call for shared protocols and independent audits: without sufficient access, it is difficult to verify either the manufacturers’ claims or allegations of ineffectiveness.

For anyone receiving a report, a paper or an editorial piece, the result may be to regard this labelling as evidence of deception. But that would be a hasty conclusion, as we have seen.

The revision history, the sources used and the ability to justify conclusions are the most helpful in reconstructing the (human) work carried out. An automated indicator, on its own, is not sufficient to refute ‘artificial’ behaviour. Nor is it sufficient to prove human authenticity.

In short, these tagging tools are supposed to increase transparency amidst the sheer chaos that is AI-generated writing, but they, in turn, introduce further confusion. We’ll just have to learn to live with it.

Copyright reserved ©

Brand connect

Loading...

Newsletter

Notizie e approfondimenti sugli avvenimenti politici, economici e finanziari.

Iscriviti