An "AI Crisis" Triggered by a Photocopier
In July 2024, Nature published an alarming study[3]:
When AI models are repeatedly trained on data generated by themselves, they gradually "go mad"—outputs become monotonous, hallucinations increase, and long-tail knowledge is forgotten.
Researchers call this "Model Collapse" or "Model Autophagy Disorder (MAD)".
Imagine a scenario: You use a photocopier to copy a photo, then use the copy to copy again, repeating this 10 times. The final image details are blurred, noise is amplified, and colors are distorted.
This is the essence of AI model autophagy: Each generation is "eating" the output of the previous generation, and quality degrades generation by generation.
Why Does It Happen? Three Fatal Cycles
1. Style Convergence: Becoming More "Like Itself"
The model's "catchphrases" and habitual patterns are copied and amplified in synthetic data. Cambridge University research found that after just 5 generations of recursive training, the model's output entropy (diversity index) dropped by 60%+. The model eventually loses diversity, like a repeater.
2. Long-Tail Forgetting: Rare Knowledge Disappears
In the real world, common knowledge accounts for 80%, and long-tail knowledge accounts for 20%. After multi-generation synthetic data training, common knowledge is amplified to 99%, while long-tail knowledge almost disappears. This causes the model to be fluent on "hot topics" but completely fail in "niche scenarios."
3. Error Loop: Bug "Heredity"
After the initial 1% hallucination content is treated as "training data," the hallucination rate may rise to 3% or even higher. Erroneous "factual statements" are regarded as "high-confidence truths" by the model after multiple rounds of self-training.
Industry Status and Limitations
Giants like OpenAI and Anthropic respond through Reinforcement Learning from Human Feedback (RLHF) and heterogeneous data mixing, but this faces high costs and scale limits.
Core Question: How to continuously ensure freshness and authenticity while scaling synthetic data?
Mentiss's Breaking Solution: Werewolf Naturally Prevents MAD
1. Zero-Sum Game = Built-in Quality Verification
In Werewolf, winning or losing is the only objective standard. If the "persuasive speech" generated by AI is not persuasive, it will lose the game immediately. This is the most honest feedback signal; low-quality data is naturally eliminated, while victory brought by high-quality strategies reinforces high-quality data.
2. Multi-Model Heterogeneous Games = Breaking Homogeneity
Our every game involves GPT-4, Claude, self-developed models, and human players. Differences in models offset each other, exposing blind spots, ensuring healthy data distribution.
3. Human Players = Continuous "Living Water" Injection
Human players' language, strategies, and on-the-spot reactions are the most valuable real anchors. Human participation breaks fixed patterns, introduces unexpected strategies, and maintains the freshness of data distribution. We keep rounds with human-AI disagreement as high-value samples and always maintain a certain proportion of real human participation.
Engineering Governance Scheme
1. Data Provenance and Quota Control
We divide data into four buckets: Human Baseline (Human Only), High-Value Mixed (Human-AI Mixed), Heterogeneous Mixed (Multi-model), and Single Model (Single Model). Strictly control the proportion of each bucket; once a single source accounts for too high a proportion, an alarm is triggered.
2. Quality Filtering Four-Layer Inspection
- Layer 1: Causal Consistency Check.
- Layer 2: Semantic Deduplication, preventing "cliché" proliferation.
- Layer 3: Style Balancing, ensuring strategic diversity.
- Layer 4: Preferential sampling of human-AI disagreement samples.
3. Sentinel Indicator Monitoring
Continuously monitor output diversity, long-tail recall rate, and hallucination rate on the pure human holdout set. Once indicators are abnormal, immediately roll back or introduce more real person games.
Conclusion: A Data Ecosystem Co-evolving with AI
The MAD problem essentially reflects a deep philosophy: AI cannot just learn itself in the "mirror," it must continuously contact the "window" of the real world.
Werewolf is exactly this unique window. It comes with built-in verification, natural heterogeneity, and continuous freshness. When we can solve the MAD problem through engineering, synthetic data is no longer a "toxic shortcut," but a "sustainable flywheel" to AGI.