1/4
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Arm’s race for data
AI models benefit from having more, high quality training data
“Scale is all you need”
most prized: high-quality information, such as published books and articles which have been written and edited by professionals
GPT-3 trained on about 300 billion “tokens”
current models: more than 10 trillion
OpenAI Admits
“Because copyright today covers virtually every sort of human expression including blog posts, photographs, forum posts, scraps of software code and government documents — it would be impossible to train today’s leading AI models without using copyrighted materials.” - OpenAI writes
Fair Use
AI companies argue their training is “fair use” or allowed under copyright law because they transformed the works for a different purpose
Fair use allows limited use of copyrighted materials without permission of the owners based on 4 non-exclusive statutory factors:
purpose and character of the use
nature of copyrighted work
amount of work copied
the use’s effect on the existing and potential market
no set criteria more case-by-case basis
Model Collapse
model collapse: a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors
this makes training data uncontaminated by AI particularly valuable
e.g. pre 2022 books
Discussion Activity
Authors Guild vs OpenAI
Bartz vs Anthropic → class action lawsuit of anthropic downloading millions of books on pirated websites for training of AI models
Kadrey vs Meta