Skip to main content
News Directory 3
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
Menu
  • Business
  • Entertainment
  • Health
  • News
  • Sports
  • Tech
  • World
AI & Copyright: New Dataset Challenges Tech Firms - News Directory 3

AI & Copyright: New Dataset Challenges Tech Firms

June 8, 2025 Catherine Williams Tech
News Context
At a glance
  • A group of AI researchers ⁣is challenging the notion that copyrighted material is essential for training artificial intelligence.
  • The researchers then used this dataset to train a 7 billion parameter language ⁢model.
  • Stella Biderman,executive director of Eleuther AI,emphasized the challenges of creating such datasets.
Original source: slashdot.org

AI firms, take ‍note: researchers are proving ‍that AI models can be trained effectively without infringing on⁢ copyright. This groundbreaking study reveals an ⁣eight-terabyte dataset built entirely from openly licensed text, challenging the industry’s persistent reliance on copyrighted material. The resulting AI model achieved performance levels comparable to industry standards, marking a potent shift toward ⁣ethical⁢ AI and open-source solutions. This initiative highlights the meticulousness required, including manual‍ annotation and thorough checking, when creating ethical datasets. Find insight into⁢ the hurdles of inconsistent ⁢data, legal challenges of licensing, and the value of datasets like those at the Library of Congress. News Directory 3 ⁤understands the ⁤importance of these innovations. Discover what’s ⁤next in the quest for clarity and ethical practices within the AI landscape.

Key Points

  • AI firms argue copyrighted material‍ is needed for training.
  • Researchers built an eight-terabyte dataset using only openly licensed text.
  • The resulting⁤ AI model performed comparably to industry ‍standards.
  • Manual annotation and checking are crucial for ethical data set creation.
  • Openness in AI training data is vital for ‍social and scientific ‍value.

Researchers Explore ethical AI Training⁢ with Open Source Data

⁤ ⁣ ⁤ updated June 7, 2025
⁢

A group of AI researchers ⁣is challenging the notion that copyrighted material is essential for training artificial intelligence. They successfully built a⁢ massive, eight-terabyte dataset using only openly licensed or public domain text.⁣ This initiative explores ethical AI and offers⁤ a⁢ potential alternative to the industry’s reliance on copyrighted sources.

The researchers then used this dataset to train a 7 billion parameter language ⁢model. The model’s⁢ performance was ‍comparable to industry⁢ efforts like ‍Meta’s Llama 2-7B, demonstrating the viability of⁢ using ethically sourced data for AI training.

Stella Biderman,executive director of Eleuther AI,emphasized the challenges of creating such datasets. The⁣ process is painstaking, tough to fully automate, and requires significant human⁤ involvement. ⁢Technical hurdles include ‍inconsistent data formatting, while legal challenges ⁤involve determining the correct ⁤licenses for various websites. The issue of improperly licensed data⁤ is widespread, Biderman noted.

Despite the difficulties, the group discovered new, ethically usable datasets, including⁣ a ⁤collection ⁤of 130,000 English language books from the Library of Congress. This collection nearly doubles the size of Project Gutenberg’s‍ popular books ⁢dataset. The⁣ initiative builds upon other efforts to develop more ethical datasets, ⁣such as FineWeb from Hugging Face.

Biderman remains skeptical that ‍this approach can generate enough content to match⁤ the ⁤scale of today’s most‍ advanced models. However, she hopes their work will⁢ encourage companies⁣ like OpenAI ⁣and Anthropic to increase transparency⁢ regarding their training⁢ data.

⁢ “Even partial transparency has ⁣a huge amount of social value and a moderate amount of⁢ scientific⁢ value,” Biderman said.

What’s next

The researchers hope their work will encourage greater transparency in the AI industry regarding training data. They also aim ‍to inspire further exploration of ethical AI training methods and ⁣the progress of high-quality, openly⁢ licensed datasets.

Share this:

  • Share on Facebook (Opens in new window) Facebook
  • Share on X (Opens in new window) X

More on this

  • Thorsten Hoppe Discovers Leucine Boosts Mitochondrial Energy Production
  • iPhone 17 vs iPhone 17e: Which One Should You Buy?

Related

Search:

News Directory 3

News Directory 3 catalogs US newspapers, news services, newsstands and digital news outlets across all 50 states. Browse local publishers by city, state, or topic, and follow current headlines linked back to their original sources.

Quick Links

  • Disclaimer
  • Terms and Conditions
  • About Us
  • Advertising Policy
  • Contact Us
  • Cookie Policy
  • Editorial Guidelines
  • Privacy Policy

Browse by State

  • Alabama
  • Alaska
  • Arizona
  • Arkansas
  • California
  • Colorado

© 2026 News Directory 3. All rights reserved.
For contact, advertising, copyright, issues email: office@newsdirectory3.com