The Nature paper published in 2024 proved something that most people in AI did not want to hear. Shumailov, Shumaylov, Zhao, Papernot, Anderson, and Gal showed that AI models trained on AI-generated data get worse. Not incrementally worse. Catastrophically worse. They called it model collapse.
The mechanism is simple. The internet is filling with AI-generated content. Blog posts, articles, reviews, comments, social media. AI companies scrape the internet to train the next generation of models. That means the next generation trains on the output of the current generation. Each cycle loses information. The rare, unusual, creative parts disappear first. The terms paper uses is tails of the distribution. The weird ideas and unexpected perspectives instead of a weighted average of everyone.
Those disappear first. What remains is the average, the safe, the expected, the bland. Then the next generation trains on that and loses more. The researchers proved this is inevitable. Even under ideal conditions. Even when some of the original human data is preserved. The output converges toward a narrow, flattened version of reality that looks nothing like the original distribution.
The scale of this is already measurable. Navtoor on X posted a thread recently that resonated because people already feel it. ChatGPT feels dumber than it used to. Prompts that worked six months ago produce worse results. The writing sounds flatter. The ideas sound safer. The internet itself feels like it is shrinking.
That feeling is not nostalgia. It is the observable effect of model collapse in progress.
The practical consequence for anyone who writes, edits, or publishes is direct. The value of human-generated content is not going down. It is going up. Every model that trains on synthetic data loses fidelity to the original human distribution. The models get blander. The internet gets blander. But the content that was produced by a human with a specific point of view becomes more scarce.
Scarcity creates value.
The companies that understand this are doing the opposite of what the market expects. They are not rushing to automate more. They are creating systems that preserve and reward human output. They are paying writers to produce work that no model would generate. They are building pipelines where AI handles the mechanical parts and humans handle the judgment parts.
The Nature paper ends with a line that should be framed in every newsroom. Access to the original data distribution is essential. In learning tasks where the tails of the underlying distribution matter, you need access to real human-produced data.
That is what this site is built on. Human-produced data. Arguments with edge. Writing that picks a side.
Model collapse is not a prediction about the future. It is a description of the present. The internet is being diluted. The only durable answer is to keep producing work that cannot be averaged.