Publication
arXiv
Stage
Preprint
What we read
Summary of the paper
Authors
Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang, Luke Zettlemoyer
Universities and research institutions
Not yet supplied in verified metadata; the Brief does not guess.

What they did and found

We compared two types: dense transformers and models made of many tiny parts. When the same data were shown repeatedly, the many-parts models lost accuracy faster, especially as sparsity grew. Some regularization helped, but not fully.

Why it matters

Practitioners should be careful reusing data with models built from many small parts. A small share of new examples plus simple masking can cut overfitting, but guarantees are not possible.

Systems built from many tiny parts split work across units. Repeating content can cause these parts to memorize data, reducing flexibility compared with a single dense model.

Example: a government portal that reuses training data should mix in new, representative content and use simple masking to limit overfitting.

What remains uncertain

Results depend on exact data mix and settings; other mixture-of-experts designs may behave differently in other tasks or with different training budgets.

Read the paper PDF ↗

Original sources · 1
  1. Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data ↗arXiv · 2026-09-10

Check the original paper for its authors, methods, version and access terms.