Cut, copy: Why books are being destroyed to feed AI, Pg12
AI companies are controversially destroying physical books through 'destructive scanning' for training data, raising significant copyright and ethical concerns among booksellers and authors.
AI companies, including Anthropic, are reportedly making bulk purchases of physical books in Australia and Europe.
These books are being subjected to "destructive scanning," a process where they are destroyed after their content is digitized for training AI models.
This practice, revealed in unsealed documents from a copyright lawsuit against Anthropic, has raised significant concerns among booksellers and copyright holders.
The demand for physical books is driven by the need for high-quality, human-written data for Large Language Models (LLMs), particularly older content free of AI-generated text.
Detailed Insights:
The destructive scanning process involves removing book bindings and feeding loose pages into automated scanners, with remaining pages often recycled or shredded.
This method is favored for its speed and ability to produce flat, high-quality images for Optical Character Recognition (OCR) compared to non-destructive scanning.
Anthropic's Project Panama was previously identified as a large-scale effort to acquire and digitize millions of physical books for building training datasets.
AI developers view this practice as a "middle path" to obtain data, balancing the need for content with avoiding contentious web scraping and negotiating licenses.
The legal implications under Indian copyright law are being debated, specifically regarding the applicability of the Fair Dealing Exception under Section 52 of the Copyright Act of 1957.
While digitisation for AI training might be considered "non-expressive uses," concerns of authors and publishers extend beyond current copyright interpretations.
Proposed solutions to address these concerns include mandatory labelling of AI-generated content and establishing funds to support authors and artists affected by AI adoption.
This current trend echoes the earlier Google Books project, which also involved mass digitisation but primarily aimed at making books searchable rather than solely for AI training.
Institutions like the Internet Archive advocate for and utilize non-destructive digitisation methods to preserve the integrity of original physical copies.
Key Concepts Involved:
Large Language Models (LLMs): AI models that learn statistical patterns from vast text data to generate human-like text.
Generative AI: A type of artificial intelligence capable of producing new content, such as text, images, or audio.
Destructive Scanning: A digitisation method that disassembles physical books by cutting their spines to facilitate high-speed scanning.
Fair Dealing Exception: A provision in copyright law, such as Section 52 of the Copyright Act of 1957, allowing limited use of copyrighted material without permission for specific purposes.