AI Training Data Set: Millions of Personal Data Examples
AI Training Data Leaks expose Millions to Identity Theft adn Privacy Violations
Table of Contents
A massive dataset used to train cutting-edge AI models, DataComp CommonPool, has been found to contain a staggering amount of sensitive personal information, raising serious privacy concerns for millions of individuals. Researchers discovered that the dataset, which comprises 12.8 billion image-text pairs, includes resumes, government identification documents, and othre personally identifiable information (PII) scraped from the public internet.
The Unsettling Contents of CommonPool
The DataComp commonpool dataset, released in 2023, was intended for academic research and the training of generative text-to-image models. though, a recent analysis revealed that a notable portion of the data is far from benign. Researchers found that many documents within the dataset were linked to real people through platforms like LinkedIn and other web searches. While some documents were validated, many others could not be due to issues like poor image clarity.
The sensitive information disclosed in these documents is extensive and alarming. It includes:
Disability status
Background check results
Birth dates and birthplaces of dependents
Race
When résumés were successfully linked to individuals with an online presence,researchers also uncovered:
Contact information
Government identifiers
Sociodemographic information
Face photographs
Home addresses
Contact information of references
Examples of the types of identity-related documents found in CommonPool’s dataset include credit cards,Social Security numbers,and driver’s licenses. While the researchers redacted personal information and paraphrased text to protect privacy in their reporting, the presence of such documents highlights a critical vulnerability.
A Legacy of Data Scraping and Potential for Widespread harm
DataComp CommonPool is not an isolated incident; it is a successor to the LAION-5B dataset, which has been instrumental in training popular AI models like Stable Diffusion and Midjourney. Both datasets draw their data from web scraping conducted by the nonprofit Common Crawl between 2014 and 2022. This shared data source means that the datasets are highly similar, and the same personally identifiable information likely exists in LAION-5B and other downstream models trained on CommonPool data.
The implications of this overlap are significant. While commercial AI models often do not disclose their training data, the common origins of CommonPool and LAION-5B suggest a widespread risk. Rachel Hong, a PhD student in computer science at the University of Washington and the lead author of the paper detailing these findings, noted that DataComp commonpool has been downloaded over 2 million times in the past two years. This widespread adoption means that “there are many downstream models that are all trained on this exact data set,” possibly duplicating these privacy risks across numerous AI applications.
Good Intentions Are Not Enough
The findings underscore a fundamental challenge in the development of large-scale AI models: the inherent difficulty in ensuring data purity and privacy when scraping the vast expanse of the internet.
“You can assume that any large-scale web-scraped data always contains content that shouldn’t be there,” states Abeba Birhane, a cognitive scientist and tech ethicist leading Trinity College Dublin’s AI accountability Lab. Her research into LAION-5B has previously identified not only personally identifiable information but also child sexual abuse imagery and hate speech within such datasets.
The creation of datasets like CommonPool, while potentially driven by the goal of advancing AI research, highlights the critical need for more robust data governance, ethical sourcing practices, and obvious data handling within the AI industry. Without these safeguards, the pursuit of AI innovation risks compromising the privacy and security of millions.
