AI Web Scraping: Impact on Science & Journals
- Scientific databases and academic journals are struggling to cope with a massive influx of automated web-scraping bots.
- DiscoverLife, an online image repository with nearly 3 million species photographs, experienced this firsthand.
- The surge in bot activity has intensified since the release of DeepSeek, a Chinese large language model.
AI bots are overwhelming scientific databases and academic journals, causing service disruptions across the board. This surge, directly linked to the rise of efficient AI models, has rendered numerous sites unusable, impacting access to crucial scientific details. DiscoverLife, a repository, saw millions of daily hits, while medical journal publisher BMJ faced bot traffic exceeding legitimate user activity. Over 90% of the Confederation of Open Access Repositories members report AI bot scraping. Admins are fighting back with rate limiting and refined detection. News Directory 3 has been keeping an eye on this story. Learn about the technological and policy changes coming to balance AI growth and scientific information accessibility. Discover what’s next in this evolving landscape.
AI Bots Overload Scientific Databases, Causing Disruptions
Scientific databases and academic journals are struggling to cope with a massive influx of automated web-scraping bots. These bots, designed to gather training data for AI models, are generating traffic volumes that render many sites unusable.
DiscoverLife, an online image repository with nearly 3 million species photographs, experienced this firsthand. Starting in February,the site was bombarded with millions of daily hits,slowing it down to the point of inaccessibility.
The surge in bot activity has intensified since the release of DeepSeek, a Chinese large language model. DeepSeek demonstrated that effective AI could be built with fewer computational resources than previously believed. This revelation triggered an ”explosion of bots,” according to industry observers, all seeking data to train similar models. The rise in AI bot traffic is impacting accessibility to important scientific resources.
The Confederation of Open Access Repositories reports that over 90% of its 66 surveyed members have experienced AI bot scraping. Approximately two-thirds of these members have suffered service disruptions as an inevitable result.The medical journal publisher BMJ has even seen bot traffic surpass legitimate user activity,overloading servers and interrupting customer services.
What’s next
Database administrators are exploring various methods to mitigate the impact of these bots, including rate limiting and more refined bot detection techniques. The long-term solution likely involves a combination of technological and policy changes to balance the needs of AI development with the accessibility of scientific information.
