Az OpenAI bejelentette, hogy az Európai Unió mesterséges intelligenciáról szóló jogszabályának megfelelve láthatatlan vízjelekkel látja el a ChatGPT és a Codex által generált szövegeket az EU-ban. A textGrain nevű technológia észrevehetetlen statisztikai mintázatot kever a szóválasztásokba, amelyet egy saját detektorszoftverrel lehet azonosítani. Az API-felhasználók világszerte választható funkcióként már most bekapcsolhatják a vízjelezést.
A rendszernek komoly korlátai vannak, a szövegek szerkesztése ugyanis jelentősen gyengíti a jelölést. Ha a szövegnek mindössze 10 százalékát szinonimákra cserélik, a felismerési arány 92 százékról 66 százalékra csökken, míg a rövidebb vagy kötöttebb szóhasználatú, például matematikai szövegeknél még nehezebb a detektálás. A vízjel nem azonosítja a felhasználót, nem igazolja a pontosságot, és nem méri az emberi közreműködés mértékét sem.
A vízjelek azonosítására szolgáló detektort az OpenAI egyelőre nem teszi nyilvánossá a téves riasztások kockázata miatt. A hozzáférést első körben kizárólag jóváhagyott kutatóknak és szakértő szervezeteknek biztosítják, akik segítenek a technológia fejlesztésében. A cég a későbbiekben nyílt forráskódúvá tervezi tenni a megoldást.
Az eredeti szöveg (OpenAI)
Promoting transparency within the limits of today’s technology.
Content provenance helps people understand where content came from, how it was created or edited, and whether it contains signals associated with our models. We’ve already made tools publicly available to identify images and audio generated by our models. Today, we’re sharing our approach to text watermarking in response to the EU AI Act, and how it fits into our broader work.
The EU AI Act requires generative AI providers to make generated text identifiable in a machine-readable way. Text watermarking and detection remain early technologies with significant limitations, and views about their benefits and responsible uses are still developing. Our phased approach reflects both the EU AI Act requirements as well as the technology’s limitations, with an emphasis on transparency about what a text watermark can and cannot tell people:
Starting today, API customers globally will be able to opt in to text watermarking for select models. Text watermarking will remain off by default in the API.
Over the coming weeks, we will add an invisible watermark to eligible ChatGPT and Codex text output in the European Union.
We’re opening applications to access our text watermark detector. Access will initially be limited to approved researchers and expert organizations that can help us evaluate and improve the technology.
The above only applies to text provenance—our verification tools for audio and images, including our openai.com/verify web tool and our Content Provenance API(opens in a new window), will continue to be publicly accessible to organizations looking to understand whether an image or audio file was generated by one of our systems.
Our text watermarking technology, textGrain, adds an invisible statistical signal to the model’s word choices. Our detector looks for that signal to assess whether a passage contains an OpenAI watermark. More details about how textGrain works can be found in our technical report(opens in a new window), which will be updated with additional details in the coming weeks. We also plan to make the technology available in open source so that others can build on it.
In our evaluations, textGrain matched or exceeded the performance of other approaches we tested, including SynthID for text. Even so, strong performance under ideal conditions does not guarantee reliable detection in everyday use.
Detectors can make two kinds of errors: they can report a watermark where none is present—a false positive—or miss a watermark that is present—a false negative. Our evaluations below illustrate some of the challenges:
Shorter or more constrained text is harder to detect. At a target false positive rate of 1%, our detector identified watermarks in about 80% of 200-token passages, compared with about 95% of 400-token passages, for content such as psychology. Detection rates were substantially lower for content such as mathematics, where there is less flexibility in word choice.
Editing can weaken the watermark. In an evaluation of 400-token passages, replacing 10% of words with synonyms reduced detection from about 92% to 66%. Replacing 25% of words reduced it to 17%.
These limitations contribute to our decision to provide initial detector access only to approved researchers and expert organizations, who can help us evaluate reliability and responsible uses.
This chart shows results for watermarked responses to mathematics and psychology questions from the ELI5 dataset(opens in a new window) at a target false positive rate of 1%. Detection improves with text length, but is substantially lower overall for content where there is less flexibility in word choice, such as mathematics.
Editing can substantially weaken the watermark signal. This chart shows how replacing 10% or 25% of the words in a passage affects detection. Results are based on watermarked English responses to questions from ELI5(opens in a new window).
Across the benchmarks we