The internet as we know it is changing. The once open web, where information flowed freely, is becoming more restricted as a growing battle between websites and AI crawlers escalates. As highlighted in a recent MIT Technology Review article, this conflict is not just about AI companies scraping data—it’s about the very future of an open and accessible web.
For years, web crawlers have been the backbone of the internet, powering search engines, price comparisons, security monitoring, and digital archives. These bots index and collect web data, making information discoverable and accessible. But the explosion of artificial intelligence has turned this long-standing relationship upside down. AI companies, including OpenAI, Anthropic, and Google, use these crawlers to harvest massive amounts of data—Wikipedia pages, news articles, Reddit discussions, and more—to train their language models. But this has triggered alarm among content creators, who fear that AI systems built on their work will ultimately replace them.
News publishers worry that AI chatbots will replace traditional readers. Artists and designers fear AI image generators will capture their clients. Developers worry that AI code generators will substitute their contributions. Websites have started to fight back, and the tension is growing. Legal and legislative actions are underway, including lawsuits and bills like the EU AI Act, but the legal process is slow. Meanwhile, websites have begun implementing restrictions to block AI crawlers, resulting in a tug-of-war that threatens to undermine the open web.
Since mid-2023, more than 25% of high-quality data has been blocked to crawlers, but AI companies, such as OpenAI, have been accused of ignoring these restrictions, leading to a series of covert battles. To mitigate the impact, web infrastructure companies like Cloudflare have started offering tools to block crawlers, which could eventually create a fragmented web, where access to data depends on who can afford the most advanced tools or the best legal team.
As the battle intensifies, large tech companies can survive by licensing data or creating advanced crawlers, but small content creators—artists, bloggers, and independent researchers—are at risk of losing their place on the web. These creators may resort to paywalls or removing their content altogether, which would diminish the web’s diversity and accessibility.
In the long term, the commercial interests of large companies may result in a web divided into exclusive territories, where only a few major players control the data. This would not only stifle competition but also reduce the richness and openness of the internet. To avoid this outcome, advocates for an open web must push for laws, policies, and technical solutions that protect non-AI uses of web data while still safeguarding the rights of content creators. The battle over crawlers is just beginning, and the stakes couldn’t be higher.

