The open internet ecosystem relies on web crawlers—automated bots that systematically browse millions of websites to collect various forms of data, including text, tables, images, audio, and video. Web-crawled data serve multiple purposes, such as:
Estimates suggest that crawler traffic accounts for nearly half of all internet activity and is poised to surpass human-driven traffic. However, the rapid expansion of AI-powered crawlers threatens the transparency and accessibility of the internet. Many websites risk displacement as AI crawlers increasingly dominate web traffic.
In response, website owners are implementing protective measures such as logins, paywalls, and anti-crawling technologies that detect, restrict, block, or charge fees for nonhuman traffic. These actions are fragmenting the internet, creating areas where AI crawlers have limited, slower, or no access—ultimately reducing information availability for human users and reshaping the concept of an "open" internet.
Large tech companies can afford to license extensive datasets and develop advanced AI web crawlers capable of circumventing these restrictions. In contrast, smaller content creators—such as visual artists, YouTube creators, and independent bloggers—may choose to hide their work behind logins and paywalls or remove it from the internet altogether. This shift risks concentrating control over the information ecosystem in the hands of AI developers and large data publishers.
To preserve an open internet, advocates will likely turn to laws, policies, and technical infrastructure aimed at protecting non-commercial and noncompetitive uses of web data.
Have something to add or refine? Your input in this work matters greatly and we look forward to reviewing your additions
Click on a star to rate it!