Here are the biggest misconceptions about AI content scraping

AI bots scraping publishers’ sites for real-time information are now scraping publishers’ sites more than the bots used to train large language models. And they’re harder to detect.

That’s according to the latest report from TollBit, a data marketplace for publishers and AI companies. From Q4 2024 to Q1 2025, bot scrapes used for Retrieval Augmented Generation, or RAG, per site grew 49%. That is nearly 2.5 times the rate of training bot scrapes (which grew by 18%) in the same time period. 

An increase in bots scraping content from publishers’ sites represents a threat to their businesses. But scraping for AI training and scraping for real-time outputs present different challenges — and some opportunities — for publishers. And not all of them are fully understood. 

Continue reading this article on digiday.com. Sign up for Digiday newsletters to get the latest on media, marketing and the future of TV.

,Read More

How Future is using its own AI engine to turn deeper engagement into ad dollars 

Future is betting on AI to boost recirculation – and make that stickier audience more appealing to advertisers.

The publisher’s new proprietary AI-powered content categorization engine, called Advisor, acts like a “brain” trained on Future’s internal data, according to Jamie Samuel, head of commercial products at Future. Built into its audience data platform Aperture, Advisor lets Future quickly test an “unlimited” number of approaches to keep users engaged on its sites, ranging from chatbots to content recommendation widgets, he said.

Advisor uses machine learning and OpenAI’s large language models to analyze articles in real-time across Future’s over 50 publications, including Marie Claire, Who What Wear, Tom’s Guide and The Week. 

Continue reading this article on digiday.com. Sign up for Digiday newsletters to get the latest on media, marketing and the future of TV.

,Read More