A watermark for chatbots can expose text written by an AI

The tool could let teachers spot plagiarism or help social media platforms fight disinformation bots.

Close up of woman's hand whiting to do list

Close up of woman’s hand whiting to do list

Hidden patterns purposely buried in AI-generated texts could help identify them as such, allowing us to tell whether the words we’re reading are written by a human or not.

These “watermarks” are invisible to the human eye but let computers detect that the text probably comes from an AI system. If embedded in large language models, they could help prevent some of the problems that these models have already caused.

For example, since OpenAI’s chatbot ChatGPT was launched in November, students have already started cheating by using it to write essays for them. News website CNET has used ChatGPT to write articles, only to have to issue corrections amid accusations of plagiarism. Building the watermarking approach into such systems before they’re released could help address such problems. 

In studies, these watermarks have already been used to identify AI-generated text with near certainty. Researchers at the University of Maryland, for example, were able to spot text created by Meta’s open-source language model, OPT-6.7B, using a detection algorithm they built. The work is described in a paper that’s yet to be peer-reviewed, and the code will be available for free around February 15. 

AI language models work by predicting and generating one word at a time. After each word, the watermarking algorithm randomly divides the language model’s vocabulary into words on a “greenlist” and a “redlist” and then prompts the model to choose words on the greenlist. 

The more greenlisted words in a passage, the more likely it is that the text was generated by a machine. Text written by a person tends to contain a more random mix of words. For example, for the word “beautiful,” the watermarking algorithm could classify the word “flower” as green and “orchid” as red. The AI model with the watermarking algorithm would be more likely to use the word “flower” than “orchid,” explains Tom Goldstein, an assistant professor at the University of Maryland, who was involved in the research. 

ChatGPT is one of a new breed of large language models that generate text so fluent it could be mistaken for human writing. These AI models regurgitate facts confidently but are notorious for spewing falsehoods and biases. To the untrained eye, it can be almost impossible to distinguish a passage written by an AI model from one written by a human. The breathtaking speed of AI development means that new, more powerful models quickly make our existing tool kit for detecting synthetic text less effective. It’s a constant race between AI developers to build new safety tools that can match the latest generation of AI models.

“Right now, it’s the Wild West,” says John Kirchenbauer, a researcher at the University of Maryland, who was involved in the watermarking work. He hopes watermarking tools might give AI-detection efforts the edge. The tool his team has developed could be adjusted to work with any AI language model that predicts the next word, he says.

The findings are both promising and timely, says Irene Solaiman, policy director at AI startup Hugging Face, who worked on studying AI output detection in her previous role as an AI researcher at OpenAI, but was not involved in this research. 

“As models are being deployed at scale, more people outside the AI community, likely without computer science training, will need to access detection methods,” says Solaiman. 

There are limitations to this new method, however. Watermarking only works if it is embedded in the large language model by its creators right from the beginning. Although OpenAI is reputedly working on methods to detect AI-generated text, including watermarks, the research remains highly secretive. The company doesn’t tend to give external parties much information about how ChatGPT works or was trained, much less access to tinker with it. OpenAI didn’t immediately respond to our request for comment. 

It’s also unclear how the new work will apply to other models besides Meta’s, such as ChatGPT, Solaiman says. The AI model the watermark was tested on is also smaller than popular models like ChatGPT. 

More testing is needed to explore different ways someone might try to fight back against watermarking methods, but the researchers say that attackers’ options are limited. “You’d have to change about half the words in a passage of text before the watermark could be removed,” says Goldstein.  

“It’s dangerous to underestimate high schoolers, so I won’t do that,” Solaiman says. “But generally the average person will likely be unable to tamper with this kind of watermark.”  

Read More
Melissa Heikkilä

Latest

Tottenham Hotspur vs Getafe

Tottenham Hotspur will return to London after their pre-season tour of Australia to take on Spanish side Getafe in a friendly at Hotspur Way on Saturday. It has been an eventful pre-season for the North London club following significant changes to a squad that struggled over the past two seasons, while Getafe arrive looking to

Bonang Matheba to receive Basadi in Music Trailblazer Award

403 ERROR Request blocked. We can't connect to the server for this app or website at this time. There might be too much traffic or a configuration error. Try again later, or contact the app or website owner. If you provide content to customers through CloudFront, you can find steps to troubleshoot and help prevent

Two-minute drill: Patrick Mahomes’ health, Bijan Robinson’s historic deal, former No. 1 overall pick’s reunion

As the Arizona Cardinals and Carolina Panthers prepare for Thursday's Hall of Fame Game in Canton, Ohio, to officially kick off the 2026 NFL preseason slate, there is a lot to digest around the league. From a three-time Super Bowl champion's health, Baker Mayfield's contract drama and new deals for Bijan Robinson and Zay Flowers

AI Framework Helps Classify Renal Tumors, Predict Outcomes

TOPLINE Researchers developed an AI framework that identified tissue components, classified nine renal cell tumor subtypes, predicted nuclear grade, and stratified survival across multicenter cohorts. The AI-derived risk score outperformed World Health Organization/International Society of Urological Pathology (WHO/ISUP) grading in predicting recurrence-free survival (RFS) and disease-specific survival (DSS...

Newsletter

Don't miss

Tottenham Hotspur vs Getafe

Tottenham Hotspur will return to London after their pre-season tour of Australia to take on Spanish side Getafe in a friendly at Hotspur Way on Saturday. It has been an eventful pre-season for the North London club following significant changes to a squad that struggled over the past two seasons, while Getafe arrive looking to

Bonang Matheba to receive Basadi in Music Trailblazer Award

403 ERROR Request blocked. We can't connect to the server for this app or website at this time. There might be too much traffic or a configuration error. Try again later, or contact the app or website owner. If you provide content to customers through CloudFront, you can find steps to troubleshoot and help prevent

Two-minute drill: Patrick Mahomes’ health, Bijan Robinson’s historic deal, former No. 1 overall pick’s reunion

As the Arizona Cardinals and Carolina Panthers prepare for Thursday's Hall of Fame Game in Canton, Ohio, to officially kick off the 2026 NFL preseason slate, there is a lot to digest around the league. From a three-time Super Bowl champion's health, Baker Mayfield's contract drama and new deals for Bijan Robinson and Zay Flowers

AI Framework Helps Classify Renal Tumors, Predict Outcomes

TOPLINE Researchers developed an AI framework that identified tissue components, classified nine renal cell tumor subtypes, predicted nuclear grade, and stratified survival across multicenter cohorts. The AI-derived risk score outperformed World Health Organization/International Society of Urological Pathology (WHO/ISUP) grading in predicting recurrence-free survival (RFS) and disease-specific survival (DSS...

Katsina unveils family planning procurement guideline to boost maternal healthcare

The Katsina State Government has unveiled domesticated guidelines for state-funded procurement of family planning commodities to strengthen healthcare financing and improve access to childbirth spacing services. The guidelines were unveiled on Wednesday at a dissemination workshop organised by the State Ministry of Health in collaboration with the Access to Medicines Initiative (AMI). The event brought

Global Business Travel Spending to Hit Record $1.71 Trillion in 2026, While Trips Reach 1.84 Billion, Says GBTA Forecast

In Brief: GBTA’s latest Business Travel Index report shows spending gains driven by higher prices as geopolitical uncertainty and transportation pressures weigh on the outlook Global Business Travel Spending to Hit Record $1.71 Trillion in 2026, While Trips Reach 1.84 Billion, Says GBTA Forecast - Image Credit Unsplash+    Global business travel spending is forecast to

Boulevard Business Park: Saudi Arabia’s Bold Bet on Mixed-Use Development

Boulevard Business Park Saudi Arabia has marked another milestone in its urban and entertainment transformation with the completion of the Kingdom’s first entertainment-focused business park and corporate resort. This remarkable achievement reflects a new approach to commercial development that thoughtfully blends business, hospitality and leisure within a single destination. More Than an Office Park Located

How Chaoshan Business Was Built on Morals, Not Contracts

This is the second article in a series on China's southern Chaoshan region, exploring the history, culture, and identities behind its recent resurgence in the spotlight. Read Part 1. In the Chinatowns of early 20th-century Bangkok, Chinese laborers who had just collected their wages would make their way to the qiaopiju, a private institution that