Skip to main content
News Directory 3
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Menu
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
AI Training Data Set: Millions of Personal Data Examples - News Directory 3

AI Training Data Set: Millions of Personal Data Examples

July 18, 2025 Lisa Park Tech
News Context
At a glance
Original source: technologyreview.com

AI Training Data Leaks expose Millions to ⁣Identity ⁣Theft⁢ adn Privacy Violations

Table of Contents

  • AI Training Data Leaks expose Millions to ⁣Identity ⁣Theft⁢ adn Privacy Violations
    • The Unsettling Contents of CommonPool
    • A Legacy of Data ⁤Scraping and Potential for Widespread harm
      • Good Intentions Are⁣ Not Enough

A massive dataset ⁣used to train cutting-edge AI models, DataComp CommonPool, has been found to contain a staggering amount of sensitive personal information, raising serious⁣ privacy concerns for millions of⁢ individuals. Researchers discovered that the dataset, which comprises ⁢12.8 billion image-text pairs, includes resumes, government identification documents, and othre personally identifiable information (PII) scraped from the public internet.

The Unsettling Contents of CommonPool

The DataComp commonpool⁢ dataset, released in 2023, was intended for academic research and the training of generative text-to-image models. though, a recent analysis revealed that a⁢ notable portion of the data is far from benign. Researchers found that ⁢many documents within the dataset were linked to real people through platforms like LinkedIn and other web searches. While some documents were validated, many others could not be due to issues like poor image clarity.

The sensitive information disclosed in these documents is extensive and alarming. It includes:

Disability status
Background check results
Birth dates and birthplaces of dependents
Race

When résumés were successfully linked to individuals with an online presence,researchers also uncovered:

Contact information
Government identifiers
Sociodemographic information
Face photographs
Home addresses
Contact information of references

Examples of the types of identity-related documents found in CommonPool’s dataset include credit⁤ cards,Social Security ⁢numbers,and driver’s licenses. While the researchers redacted personal information and paraphrased text to protect privacy in their reporting, the presence of such documents highlights a critical vulnerability.

A Legacy of Data ⁤Scraping and Potential for Widespread harm

DataComp CommonPool is not an isolated incident; it is a successor to ⁤the LAION-5B dataset, ⁢which has been instrumental in training popular AI models like Stable Diffusion and Midjourney. Both datasets draw their data from web scraping conducted by the nonprofit Common Crawl between⁤ 2014 and 2022. This shared data source means that the datasets are highly similar, and the same personally identifiable information likely exists in LAION-5B ⁤and other downstream models trained on CommonPool data.

The implications of ‍this overlap are significant. ⁣While commercial AI models often do not disclose their ⁣training data, the common ⁢origins of CommonPool and ⁤LAION-5B suggest⁤ a widespread risk. Rachel Hong, a PhD student in computer science at the University of Washington and the lead author of⁤ the paper detailing these findings, ‍noted that DataComp commonpool has been downloaded over ⁤2 million times in the past two years. This widespread adoption means that “there are many downstream models that are all trained on this exact data set,” possibly duplicating these privacy risks across numerous AI applications.

Good Intentions Are⁣ Not Enough

The findings underscore a fundamental challenge in the development of large-scale AI models: the inherent difficulty in ensuring data purity and privacy when scraping the vast expanse of the internet.

“You can assume that any large-scale web-scraped data always contains content that shouldn’t be there,” states Abeba Birhane, a⁤ cognitive scientist and ⁣tech ethicist ⁢leading Trinity College Dublin’s AI accountability Lab. Her research into LAION-5B has previously identified not only personally identifiable information but also child sexual abuse imagery and⁣ hate speech within such datasets.

The creation of datasets like CommonPool, while potentially driven by the goal of advancing AI research,⁤ highlights the critical need for more robust data governance, ethical sourcing practices, and obvious data⁢ handling within the ⁢AI ‍industry. Without these safeguards, the pursuit⁣ of ⁢AI innovation risks compromising the ⁢privacy and security of millions.

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

Worth a look

  • Android users have access to several built-in text entry methods
  • Palantir reports second-quarter revenue of 1.935 billion dollars
  • Third Circuit Rules AI Training Can Infringe Copyright (time.news)
  • Teacher Carina Gubiotti Graduates from AI Training Program (archynewsy.com)

Related

Search:

News Directory 3

News Directory 3 catalogs US newspapers, news services, newsstands and digital news outlets across all 50 states. Browse local publishers by city, state, or topic, and follow current headlines linked back to their original sources.

Quick Links

  • Disclaimer
  • Terms and Conditions
  • About Us
  • Advertising Policy
  • Contact Us
  • Cookie Policy
  • Editorial Guidelines
  • Privacy Policy

Browse by State

  • Alabama
  • Alaska
  • Arizona
  • Arkansas
  • California
  • Colorado

© 2026 News Directory 3. All rights reserved.
For contact, advertising, copyright, issues email: office@newsdirectory3.com