In March 2024, NBC News found Facebook pages posting pictures of Jesus built from shrimp, crabs, and seahorses. The captions often said a person had painted them by hand. Commenters typed Amen by the thousand. 404 Media traced those pages to ad farms and scams.
So what is going on? The dead internet theory says most of what you scroll past is made by bots, and many of the accounts liking and commenting are bots too. It started as a forum post in January 2021. People laughed at the conspiracy parts. The traffic numbers kept backing up the bot-share part.
Then platform CEOs started saying it out loud. In September 2025, Sam Altman posted on X:
Loading tweet...
Time covered the post on 10 September. Reddit co-founder Alexis Ohanian had posted in June that most of what people see online was already machine-made. Kaitlyn Tiffany wrote about the theory for The Atlantic in August 2021. Four years from a forum joke to the people who run the platforms agreeing on the count.
Your group chat is probably real. The public feed, the search box, and the server logs are another story. What hurts publishers is this: the web pays for itself when a bot visit brings a person. Ads, subscriptions, and affiliate links need someone on the page. More bots now copy the page, answer the question somewhere else, and leave the site with the hosting bill.
What the numbers say
Imperva's 2016 Bot Traffic Report put automated visits at 51.8% across 100,000 domains. The 2023 report had 47.4% for 2022. The 2024 report had 49.6% for 2023. In 2018, New York Magazine found fake YouTube views, fake followers, and stolen ad impressions. Those bots sold fake attention. The newer ones read real pages so a chatbot can answer without sending a click.
In August 2026, 35.6% of traffic on Cloudflare's network came from bots. Cloudflare already drops about 6% of global traffic as malicious, and millions of customers have turned on AI crawler blocking, so the real bot share is higher than the published number.
Three cloud providers send nearly a third of all bot traffic. Amazon at 14.7% across two AS numbers. Microsoft at 9%. Google at 8.5%. The United States sends 42.6% of all bot requests.
The source is Cloudflare Radar: 330 cities, 81 million HTTP requests per second. That is the widest public count available.
Most of this traffic is still search and SEO, not model training. The Cloudflare 2025 Year in Review found that search engine crawlers made up 40% of verified bot traffic. AI crawlers accounted for 20%. SEO bots that track rankings and backlinks made up 13%. Googlebot alone generated more than a quarter of all verified bot traffic and 4.5% of HTML requests. All other AI bots combined generated 4.2%.
User-action crawling grew 15x in 2025, faster than search, training, or SEO bots. Someone asks a chatbot a question. The chatbot fetches a page. The person reads the answer and never sees the site.
Harmless spam, or a broken deal?
Fake likes and fake followers look stupid until you follow the money. They pump ad revenue. They build follower counts that make the next post look real. The shrimp Jesus posts were silly on the surface. They also taught the feed to treat bot-made images as normal.
Crawlers work the same way, at bigger scale. In July 2025 Cloudflare published the crawl-to-refer ratio: pages crawled versus visitors sent back.
Search used to crawl you and send people back. That deal paid for most of the open web. Google crawled at roughly 3 pages per referred visitor in early 2025. DuckDuckGo stayed below 1:1. Bing ranged between 50 and 70:1.
AI platforms were nowhere near those ratios. In the July 2025 Radar post, Cloudflare put Anthropic as high as 70,900 HTML pages crawled per referred visitor. By July 2025 in a later write-up, the same firm's table had Anthropic at 286,930 to 1 in January and 38,066 to 1 in July. OpenAI peaked at 3,700 to 1 in March 2025, then fell as ChatGPT search gained usage. Perplexity sent the most visitors back of the big AI names, below 200 to 1 from September 2025 onward. The 70,900 figure is one month's number, not a fixed rate. Even the better July figure is still about 10,000 times worse than Google.
Cloudflare's CEO wrote that getting traffic from OpenAI is 750 times harder than it was from Google in the previous decade. Getting traffic from Anthropic is 30,000 times harder. On 1 July 2025, Cloudflare declared "Content Independence Day" and announced it would default to blocking AI crawlers that do not pay creators.
In July 2024, Cloudflare published its first network-wide survey of AI crawler traffic. The top 4 crawlers by request volume were Bytespider from ByteDance, Amazonbot, ClaudeBot from Anthropic, and GPTBot from OpenAI. Bytespider accessed 40.40% of Cloudflare-protected websites. GPTBot accessed 35.46%. ClaudeBot reached 11.17%. Among the top 1 million websites, 38.73% were accessed by AI bots and only 2.98% took action to block them. Among the top 10 sites by traffic, 80% were accessed and 40% blocked.
The 2025 data shifted. Googlebot's crawl volume beat every other AI bot. Anthropic's ClaudeBot crawling doubled in the first half of 2025 before declining. PerplexityBot grew 3.5x. Bytespider continued the decline it started in 2024.
Who follows the rules?
Web crawlers have been around since 1994, when Martijn Koster proposed the Robots Exclusion Protocol after his server was overwhelmed by a poorly written crawler. A website puts a robots.txt file on its server. A crawler is supposed to read it and obey. Nobody enforces it. RFC 9309, published in 2022, states explicitly that robots.txt rules "are not a form of access authorization." It is a request.
Every major AI company publishes docs for its crawlers. OpenAI lists 4: GPTBot for training, OAI-SearchBot for search, ChatGPT-User for user-driven visits, and OAI-AdsBot for ad verification. Google publishes Googlebot for search and AI training combined, plus Google-Extended as an opt-out for Gemini training. Perplexity lists PerplexityBot for search and Perplexity-User for user visits. Anthropic's documentation sits behind a support page.
The docs do not stop anyone.
In August 2025, Cloudflare published that Perplexity was spoofing a Chrome user agent, rotating IPs across AS numbers, and sometimes not fetching robots.txt. Cloudflare put the declared crawler at 20 to 25 million daily requests and the stealth one at another 3 to 6 million, then de-listed Perplexity from its verified bot directory. Perplexity said Cloudflare had mixed user visits with scrapers and blamed some of the stealth traffic on Browserbase.
OpenAI's ChatGPT-User fetched robots.txt and followed it. A crawler can obey. Some choose not to.
Googlebot is the harder problem. Cloudflare's September 2025 AI bot principles blog pointed out that Googlebot crawls for both search indexing and AI training in a single pass. A website that wants to appear in Google Search must allow Googlebot. That same crawl feeds Google's AI models. Google-Extended, the opt-out for Gemini training, explicitly does not apply to AI Overviews, which are attached to search. As Cloudflare put it, the combined crawler "forces an impossible choice onto website owners."
OpenAI splits its crawlers by purpose. A publisher can allow OAI-SearchBot and block GPTBot. Google does not offer the same split.
What happens when nobody clicks through?
Pew Research Center published data in July 2025 from 68,879 Google searches across 900 US adults. Users who saw an AI-generated summary clicked on a traditional search result 8% of the time. Users who did not see a summary clicked 15% of the time. Users clicked on a link within the AI summary itself in 1% of visits.
The same study found that 18% of Google searches produced an AI summary in March 2025. Wikipedia, YouTube, and Reddit supplied 15% of the sources inside those summaries. Government websites appeared more often in AI summaries than in standard results.
SparkToro's 2024 study put US Google searches that ended without a click at 58.5%. Mobile was worse than desktop then, and worse again later. AI Overviews launched in May 2024. Google was already answering most queries on the results page. The summaries made that cheaper for Google and worse for publishers.
In March 2024, Google said its search results were being flooded by websites that "feel like they were created for search engines instead of people." The company's blog post on its March 2024 core update promised to cut "low-quality, unoriginal content" in search results by up to 40%. The bots scraping the web were also filling it with pages built to be scraped.
When bots break the vote
In January 2026, Ohanian and Kevin Rose opened a Digg public beta. Two months later the CEO, Justin Mezzell, shut the beta and downsized the team. He wrote that they had banned tens of thousands of accounts and still could not keep bots from wrecking the vote. In May 2026 the brand came back as an AI-first news product.
You cannot run a human vote if most voters are scripts. Digg found that out the hard way.
The next bots cost more than today's
Browser agents will make text crawlers look cheap. Anthropic announced Claude computer use in October 2024, enabling Claude to view screens, move cursors, click buttons, and type. Google launched Project Mariner in December 2024, a Chrome extension that controls the browser to complete tasks. The open-source Browser Use framework lets any agent navigate websites programmatically.
A text crawl fetches 2 or 3 HTTP requests. A browser agent loads every image, every JavaScript bundle, every tracking pixel, every analytics call. It renders the page, screenshots it, and feeds the image to a vision model. One agent session can hit hundreds of servers. That can equal 100 ordinary crawler visits.
A crawler at least indexes you. A browser agent uses the page to book a flight, file a ticket, or check a price, then leaves. Nobody saw your brand. Nobody subscribed. Nobody hit an ad. You still paid to serve the page.
Who pays for the pages?
A website that serves a billion pages to humans makes money from the humans who arrive. A website that serves a billion pages to AI crawlers only gets a bigger hosting bill.
Ads and subscriptions only work if a person shows up. A crawler does not click, subscribe, watch an ad, or come back. Crawl traffic used to be cheap enough to ignore. A 70,900-to-1 ratio is not. That site is paying Anthropic's bandwidth out of its own hosting invoice.
The Cloudflare Pay Per Crawl marketplace, launched in July 2025, lets publishers set a price per 1,000 tokens for crawler access. It checks who is crawling using verified bot IDs. Whether it works depends on whether AI companies find paying cheaper than getting sued.
The New York Times sued OpenAI and Microsoft in December 2023 over using its articles to train models. On 30 April 2024, eight Alden papers including the Chicago Tribune and the Denver Post filed their own suit in the same district. Those cases ask whether copyright covers the crawl. No court ruling yet. Licensing will cover the Times and a handful of other names. Forums, blogs, and local papers will not get a deal. They have no lawyers.
Common Crawl is a non-profit that keeps a free, open dump of web crawl data used for AI training. Its crawler, CCBot, obeys robots.txt and supports crawl-delay. Its July 2026 archive contains 2.14 billion web pages, 364 TiB of uncompressed content. Commercial labs train on that dump and still send almost no traffic back.
The IETF made robots.txt official as RFC 9309 in September 2022. In April 2026, the IETF AI Preferences Working Group published draft-ietf-aipref-vocab-06, with separate flags for training and search. The draft is a first try. Cloudflare shipped blocking tools and Pay Per Crawl years before the standard landed. Robots.txt was a rule any site could post. Pay Per Crawl only works if you sit on Cloudflare.
If crawl traffic is free to take and expensive to serve, sites lock it down. Paywalls, logins, bot checks, apps. The companies that wanted an open crawl are giving sites a reason to close. The sites that stayed open the longest pay the most for staying open.
A site used to bet that a visit meant a reader who might pay. Most visits are bots now. The publisher still hosts the page. The bot copies it and sends nobody back. Suits and new standards have not paid the bill.

