

Google already did this ~15 years ago with the google library project, but they didn’t buy books. They took them from libraries and then as a result of scanning old and rare books, they were generally damaged or destroyed.
I know. I saw it first hand. There wasn’t an incinerator or anything, but the carts full of books that disintegrated half the time. They had quotas for pages scanned and falling below them was a firable offense. You had to rush through it and there was no process to set aside books that disintegrated when the pages were turned.








I promise google used the scanned books with OCR to train their LLMs. This sort of thing isn’t even new from them.