Tech · Apple Machine Learning
As language models scale, the amount of data they require scales up, yet many target data sources
Compiled by KHAO Editorial — aggregated from 1 source. See llms.txt for citation guidance.
★ Tier-1 Source
Apple study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific, and quality-filtered mixtures.
Key facts
- Authors Anastasiia Sedova, Skyler Seto, Natalie Schluter, Pierre Ablin
- Scaling Laws for Mixture Pretraining Under Data Constraints
- The team study this trade-off across more than 2,000 language-model training runs spanning multiple model and target dataset sizes, as well as several data types, including multilingual, domain-specific
- Next, they introduce a repetition-aware mixture scaling law that accounts for the decreasing value of repeated target tokens and the regularizing role of generic data
Summary
Authors Anastasiia Sedova, Skyler Seto, Natalie Schluter, Pierre Ablin. As language models scale, the amount of data they require grows, yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size.