Models with many small parts overfit on repeated data compared to dense models
Data repetition hurts sparse models even as some safeguards help.
- Publication
- arXiv
- Stage
- Preprint
- What we read
- Summary of the paper
- Authors
- Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
- Universities and research institutions
- Not yet supplied in verified metadata; the Brief does not guess.
What they did and found
We compared two types: dense transformers and models made of many tiny parts. When the same data were shown repeatedly, the many-parts models lost accuracy faster, especially as sparsity grew. Some regularization helped, but not fully.
Why it matters
Practitioners should be careful reusing data with models built from many small parts. A small share of new examples plus simple masking can cut overfitting, but guarantees are not possible.
Systems built from many tiny parts split work across units. Repeating content can cause these parts to memorize data, reducing flexibility compared with a single dense model.
Example: a government portal that reuses training data should mix in new, representative content and use simple masking to limit overfitting.
What remains uncertain
Results depend on exact data mix and settings; other mixture-of-experts designs may behave differently in other tasks or with different training budgets.
Original sources · 1
- Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data ↗arXiv · 2026-09-10
Check the original paper for its authors, methods, version and access terms.
What is your take?
Ask a question, add useful context or share a different perspective. Keep the conversation respectful and grounded.