Scaling
1 item across 1 edition. First seen Sat 12 Sep, last seen Sat 12 Sep.
- The paper, posted to arXiv on 10 September 2026, reports that "MoEs degrade more rapidly under data repetition", with the effect growing as sparsity increases: dense 80M-parameter models tolerate "8x" repetition with minimal decline while MoEs "begin to suffer at 4x" and underperform dense alternatives at "32x".
- The study spans models from "80M to 1B active (8.5B total) parameters". The authors are Atindra Jha, Margaret Li, Jure Leskovec, Percy Liang and Luke Zettlemoyer.
- With strong masking-based regularisation, MoEs keep their advantage over dense models "even when data is repeated more than 64 times", though the paper says no method fully recovers all-unique-data performance — a direct constraint on sparse architectures as high-quality text runs short.
- Preprint, not peer reviewed. The largest configuration is 8.5B total parameters, well below frontier scale, and the paper does not claim the thresholds transfer.