Here is an update on a story you may not yet have heard of. It manages to be at once fascinating, deeply disturbing and a perfect living metaphor of our age. AI models are vulnerable to something called “model collapse.” In the early days of LLM-based AI there was a seemingly limitless store of online content to train it on. But that’s mostly been used up. Meanwhile, a rapidly growing percentage of online content is produced by AI. If you try to train AI on the product of AI it’s a bit like locking a person in a sealed room and watching them slowly suffocate as they breath in the product of their cellular respiration (CO2) rather than the fuel (oxygen) that makes the body function. Not a perfect analogy but you get the idea. This has made the creators of AI increasingly desperate in their search for new artifacts of human intellection — a phrase I heard once from the great Chaucer scholar John V. Fleming who died this past May — to feed into large language model AI systems.
The solution? Physical books published before 2022. You can guarantee as a matter of absolute certainty that whatever their quality, their merits, whatever they’re about, they are the product of human minds.
It’s relatively easy to scan a book without damaging it. But if you’re doing it at scale the costs of scanning and digitizing a book’s contents without damaging it are vastly higher. So what AI companies appear to be doing is buying up books, cutting off the spines, feeding them into industrial scale scanners and then throwing away.