❤2🤬1
Redstone and Ontology Research Unit ¦ #укртг 🧶
Усе нове це те саме старе AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
І якби ж це обмежувалося лише цим,
але
AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process.
"The optics problem is real," - ISBNdbs site says.
«AI company destroys two million books» is not a headline that generates sympathy."
"I personally have mixed feelings about all of this," - the bookseller, who suspects he's sold hundreds of books to AI companies for training data, told me.
This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
Internal Anthropic documents about it's plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process.
A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both "high volume destructive and non-destructive book scanning" services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed.
"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy," - Alsup wrote in his rulling.
This, Alsup said, was "clearly transformative" and therefore qualified as fair use under Section 107 of the Copyright Act.
ISBNdb’s site advertises this legal argument to AI companies as well.
але
AI companies’ attempts to hoover up printed books for training data got wide attention in January after a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process.
"The optics problem is real," - ISBNdbs site says.
«AI company destroys two million books» is not a headline that generates sympathy."
"I personally have mixed feelings about all of this," - the bookseller, who suspects he's sold hundreds of books to AI companies for training data, told me.
"It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped."
This bookseller said his inventory is full of rare, foreign language, and low circulation books, meaning that if they are destroyed in the process of becoming training data, they’ll be even harder to obtain.
Internal Anthropic documents about it's plan to scan millions of books, revealed in the copyright lawsuit, don’t make clear why the company wanted to destroy the books in the process.
A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both "high volume destructive and non-destructive book scanning" services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropic’s creation of digital copies of the books was legal specifically because the books were destroyed.
"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy," - Alsup wrote in his rulling.
"The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company."
This, Alsup said, was "clearly transformative" and therefore qualified as fair use under Section 107 of the Copyright Act.
ISBNdb’s site advertises this legal argument to AI companies as well.
❤2🤬2
Forwarded from remote viewers division
відверто кажучи вся стаття викликає тисяча тед-качинський.жпег реакцій
такі проекти як internet archive, інші онлайн-бібліотеки та ОСОБЛИВО піратські ресурси треба берегти й підтримувати копійкою по можливості, тому що, виявляється, тепер не можна покладатися тільки на фізичні примірники
такі проекти як internet archive, інші онлайн-бібліотеки та ОСОБЛИВО піратські ресурси треба берегти й підтримувати копійкою по можливості, тому що, виявляється, тепер не можна покладатися тільки на фізичні примірники
💯2
Redstone and Ontology Research Unit ¦ #укртг 🧶
a copyright lawsuit from book authors against Anthropic revealed internal documents detailing its plan to obtain and scan millions of printed books, and destroy them in the process
Ars Technica
Anthropic destroyed millions of print books to build its AI models
Company hired Google's book-scanning chief to cut up and digitize "all the books in the world."
❤2🤬1