Perplexity AI Uses Stealth Tactics to Bypass Google’s No-Crawl Rules
# Perplexity Accused of Using ’Stealth Bots’ to Bypass Website Restrictions
AI search engine Perplexity is facing accusations of employing deceptive tactics - including “stealth bots” – to circumvent websites’ directives against crawling, a practise that could upend long-standing internet protocols. Network security and optimization service Cloudflare publicly detailed these allegations on monday, raising concerns about the ethical implications and potential violations of established web standards.
## Cloudflare’s Examination: A Pattern of Evasion
Cloudflare researchers began investigating after receiving reports from customers who had actively blocked Perplexity’s known crawlers using standard methods. These methods included configuring their `robots.txt` files – a cornerstone of internet etiquette – and implementing rules within their Web Application Firewalls (WAFs). Despite these preventative measures, customers continued to detect Perplexity accessing their content.
The investigation revealed a concerning pattern. When confronted with blocks, Perplexity allegedly switched to a secondary, undeclared crawler designed to mask its activity. This “stealth bot” employed a range of techniques to avoid detection, effectively operating in the shadows.## Millions of Requests from Rotating IPs and ASNs
“This undeclared crawler utilized multiple IPs not listed in Perplexity’s official IP range, and would rotate through these IPs in response to the restrictive `robots.txt` policy and block from Cloudflare,” the researchers explained in a blog post. “In addition to rotating IPs, we observed requests coming from different ASNs in attempts to further evade website blocks.”
The scale of this alleged activity is significant. Cloudflare estimates the stealth crawler accessed content across “tens of thousands of domains” and generated “millions of requests per day.”
Here’s a visual portrayal of the technique Cloudflare alleges Perplexity used:
## A Challenge to Decades-Old Internet Norms
If these allegations are accurate, Perplexity’s actions represent a serious breach of internet etiquette. The foundation for controlling web crawler access was laid in 1994 with the proposal of the Robots Exclusion Protocol by engineer martijn Koster.
This protocol introduced the `robots.txt` file – a simple text file placed at the root of a website – allowing site owners to clearly communicate which parts of their site crawlers are *not* permitted to access. For nearly three decades, this system has been largely respected, forming a crucial element of how the web functions. It’s a system built on trust and mutual respect between website owners and those who index the web.
The protocol was formally standardized under the Internet Engineering Task Force in 2022, solidifying its importance in the modern internet landscape. Bypassing these directives isn’t just a technical issue; it’s a challenge to the fundamental principles of web governance.
## What Does This Mean for Website Owners and the Future of Search?
The implications of Perplexity’s alleged behavior are far-reaching. Website owners rely on `robots.txt` and WAF rules to manage server load, protect sensitive information, and control how their content is presented in search results.If crawlers can simply ignore these directives, it undermines the ability of site owners to manage their online presence.
This situation also raises questions about the ethics of AI-powered search engines. While aggressive crawling might help improve the speed and comprehensiveness of search results, it shouldn’t come at the expense of respecting website owners’ wishes and established internet protocols.
The coming weeks will likely see further scrutiny of Perplexity’s crawling practices and a broader discussion about the responsibilities of AI in navigating the web
