For millions of internet users, completing a CAPTCHA has become an almost invisible part of daily life. Clicking on images of buses, tractors or traffic lights, sliding puzzle pieces into place, or simply ticking a box marked “I’m not a robot” has long been accepted as the price of accessing websites. According to a report by El País, however, those routine interactions have once again become the focus of questions over whether the world’s largest technology company has been using them to help develop artificial intelligence systems.
The renewed scrutiny has been driven by a viral social media post and a recent change in Google’s contractual terms governing reCAPTCHA, the internet’s most widely used CAPTCHA service. The change, which reclassified Google’s role from data controller to data processor, reignited longstanding claims that information generated by users while solving CAPTCHA challenges has served purposes extending beyond website security.
CAPTCHAs were introduced around 15 years ago with two principal objectives. They were designed to distinguish human users from automated bots attempting to abuse online services while also generating large datasets that could be used to improve machine learning systems. The report notes that this dual purpose has existed since the technology’s earliest years, although questions remain over how extensively the resulting data has been applied to artificial intelligence development.
The theory circulating online is straightforward. Every time users correctly identify buses, tractors, traffic lights or other objects in a series of images, they effectively create accurately labelled datasets. Those datasets are considered highly valuable for supervised machine learning, particularly for computer vision systems that must recognise objects in real-world environments.
Karina Gibert, professor at the Universitat Politècnica de Catalunya and founder of the Intelligent Data Science and Artificial Intelligence Research Center, told El País that most CAPTCHA image databases are closely connected to traffic environments. She said such images are particularly relevant for recognising vehicles, road signs and obstacles, making them directly applicable to the technologies required for autonomous driving systems.
That connection extends to Waymo, Alphabet’s self-driving vehicle company. According to the report, accurately labelled images containing buses, tractors and traffic signals could help improve the algorithms used by autonomous vehicles operating under a wide variety of conditions.
Artificial intelligence researcher Jordi Nin explained that training computer vision systems requires large quantities of human-labelled data. He said blurred or partially obscured images become especially valuable because they help algorithms learn how humans identify objects despite imperfect conditions. Multiple users are typically shown the same images, allowing the system to confirm labels through repeated agreement before incorporating them into training datasets.
Gibert said the ability to identify traffic lights from different angles and under varying weather conditions illustrates why extensive image databases are essential for machine learning. Algorithms trained on sufficiently diverse datasets can eventually recognise objects regardless of lighting, weather or environmental conditions.
As El País reports, the origins of reCAPTCHA reveal that the technology always served purposes beyond security. The service was created by Guatemalan computer scientist Luis von Ahn, who later co-founded Duolingo. Initially, reCAPTCHA protected websites by presenting challenges that automated programs could not solve. At the same time, it asked users to identify difficult-to-read words extracted from scanned books, allowing millions of human responses to improve the digitisation of historical texts.
Google acquired reCAPTCHA in 2009 and expanded its use across the internet. The company subsequently used the service during efforts to digitise millions of books and newspapers while improving algorithms capable of reading printed text. During that period, Google openly acknowledged that CAPTCHA responses contributed to artificial intelligence development. An earlier version of the official reCAPTCHA website carried the slogan “Stop a bot, build a bot” and explained that high-quality images labelled by humans were assembled into datasets suitable for training machine learning systems.
According to Google Spain, that practice changed in 2018 with the introduction of reCAPTCHA v3. The company told El País that advances in artificial intelligence and changing cyber threats eliminated the need for visual challenges. Google said it no longer uses responses to reCAPTCHA challenges to train its models and that data collection now serves only to improve CAPTCHA performance itself. References to AI model training have also been removed from the service’s official documentation.
Even so, the report says several questions remain unresolved. Many websites continue to operate earlier versions of reCAPTCHA that relied on visual challenges. In addition, Google offers both enterprise and free versions of the service. Enterprise terms explicitly state that collected data is used only for security purposes. The free version previously operated under Google’s broader privacy policy, which states that user data may be used to improve existing services or develop new ones, leaving broader possibilities open.
The debate has also attracted academic attention. A 2023 paper from the University of California, Irvine argued that Google was using earlier versions of reCAPTCHA to obtain labelled images and AI training data. Gibert told El País that she assumes Google has used such information to train algorithms capable of recognising objects found in traffic environments, adding that supervised learning depends upon extensive collections of accurately labelled data generated by humans.
Researchers interviewed by El País agreed that high-quality training data has become one of artificial intelligence’s most valuable resources. Nin said technology companies increasingly provide access to software tools and computing resources, but the most difficult asset to obtain remains the datasets required to build advanced AI models. As long as companies retain control of those datasets, he said, competitors face significant barriers to developing systems of comparable quality.
The latest development came on 2 April, when Google formally changed its legal role within reCAPTCHA from data controller to data processor. Under the new arrangement, websites using the service become the data controllers, while Google processes information under the more restrictive terms of Google Cloud Platform agreements. Gibert said the change primarily shifts responsibility for data control to website operators while allowing Google to continue collecting and analysing information within the limits of those agreements.
The change has altered the legal framework surrounding reCAPTCHA, but questions over how user-generated data has been used, and how it may be used in the future, continue to fuel debate surrounding one of the internet’s most familiar security tools.

