AI & Copyright: New Dataset Challenges Tech Firms
- A group of AI researchers is challenging the notion that copyrighted material is essential for training artificial intelligence.
- The researchers then used this dataset to train a 7 billion parameter language model.
- Stella Biderman,executive director of Eleuther AI,emphasized the challenges of creating such datasets.
AI firms, take note: researchers are proving that AI models can be trained effectively without infringing on copyright. This groundbreaking study reveals an eight-terabyte dataset built entirely from openly licensed text, challenging the industry’s persistent reliance on copyrighted material. The resulting AI model achieved performance levels comparable to industry standards, marking a potent shift toward ethical AI and open-source solutions. This initiative highlights the meticulousness required, including manual annotation and thorough checking, when creating ethical datasets. Find insight into the hurdles of inconsistent data, legal challenges of licensing, and the value of datasets like those at the Library of Congress. News Directory 3 understands the importance of these innovations. Discover what’s next in the quest for clarity and ethical practices within the AI landscape.
Researchers Explore ethical AI Training with Open Source Data
updated June 7, 2025
A group of AI researchers is challenging the notion that copyrighted material is essential for training artificial intelligence. They successfully built a massive, eight-terabyte dataset using only openly licensed or public domain text. This initiative explores ethical AI and offers a potential alternative to the industry’s reliance on copyrighted sources.
The researchers then used this dataset to train a 7 billion parameter language model. The model’s performance was comparable to industry efforts like Meta’s Llama 2-7B, demonstrating the viability of using ethically sourced data for AI training.
Stella Biderman,executive director of Eleuther AI,emphasized the challenges of creating such datasets. The process is painstaking, tough to fully automate, and requires significant human involvement. Technical hurdles include inconsistent data formatting, while legal challenges involve determining the correct licenses for various websites. The issue of improperly licensed data is widespread, Biderman noted.
Despite the difficulties, the group discovered new, ethically usable datasets, including a collection of 130,000 English language books from the Library of Congress. This collection nearly doubles the size of Project Gutenberg’s popular books dataset. The initiative builds upon other efforts to develop more ethical datasets, such as FineWeb from Hugging Face.
Biderman remains skeptical that this approach can generate enough content to match the scale of today’s most advanced models. However, she hopes their work will encourage companies like OpenAI and Anthropic to increase transparency regarding their training data.
“Even partial transparency has a huge amount of social value and a moderate amount of scientific value,” Biderman said.
What’s next
The researchers hope their work will encourage greater transparency in the AI industry regarding training data. They also aim to inspire further exploration of ethical AI training methods and the progress of high-quality, openly licensed datasets.
