TL;DR
AI scraped it to death on behalf of prediction markets. They couldn’t handle the load.
Not only scraped, but hacked by people looking for information not-yet-public.

A year ago, Cloudflare launched pay-per-crawl, letting sites charge AI crawlers per page. Last week, it went further, announcing a pay-per-use model in which publishers get paid when their content actually appears in an AI answer, and declaring that, from 15 September, its customers’ ad-supported pages will block unpaid “mixed-use” crawlers by default.
I’d be OK paying a subscription fee to access websites.
Two BIG problems though:
a) my anonymity b) conglomeration of websites into increasingly enshittified monopolies.
It doesn’t read to me that you or I as normal people looking at a website would be paying these fees.
Correct, but if humans get in free and bots don’t, bots will pretend to be human. The only way it works ultimately is if every access is charged.
Hence the market in residential proxies (often unknown to those whose residential connections are being used by AI companies for scraping)
Capitalism, uh, finds a way.
How come none of these sites being hit by AI to the point they can’t function implement rate limiting?
Rate limiting grouped by IP or UA tends to work pretty well, and normal users have very slow request rates so generally never get impacted.
Residential proxies are a thing.
Also speaking from experience, especially asia based crawlers will send 5 requests at once from one IP, then 5 from the next etc. and the first IP wont appear again for several hours, making ip based rate limiting useless.
In my case i have honeypot links that block them, i block several data center adress ranges, have an automated whitelist for registered users, use some public blocklists and user agent filtering and that takes care of 99% of it. But it is non-trivial compared to rate limiting.
True, I also run crowdsec which blocks excessive 404s and other unusual requests too, between that and rate limiting I don’t seem to have any server overload issues for the most part.
In the article they talk about getting hit by millions of different IPs. It’s probably on a different level compared to what a random, non-targeted site has to deal with.
A few years back when graphics cards were getting botted for resale, a friend of a friend was buying with bots using a (pricey) rotating list of hundreds of proxy addresses that would each make attempts at a very reasonable looking frequency. I’m picturing something like that, except with much more funding.
I’m having a hard time blaming this entirely on Ai. We’re talking about a 30 year old site that has probably never had a proper redesign (front and back end). Of course it’s eventually going to be compromised.
Some of these used the site using legitimate URLs, others were looking for back doors, most likely so they could get to the data before it appeared on the site, or to manipulate the data presented to users.
I’m sure they know what they’re doing, making me very wrong, but bots probing for back doors is very common, and not at all anything to be concerned about. The attackers most likely don’t even care about the data, all they want is another site to spread malware or to join their botnet.
AI scrapers can be brutal though, and is very relevant.
This doesn’t say anything.
Why did the server crash? Who knows!
Did it crash because of ai scrapers? Maybe! It was definitely stressed by them.
Was the server hacked? No idea!
If it was hacked, was it an attempt to rig prediction markets? Sounds plausible!
This nebulous theorizing really is worrying me! I’m worried for the critical thinking capacity of every commenter. Sure, ai bot traffic is a problem, the entire internet is dealing with it, but its tenuous link to TheNumbers.com mystery is merely clickbait journalism without any sourcing.






