Pretraining progress is mostly coming from data

Breaking down 6 years of pretraining progress into data vs model improvements
Pretraining progress is mostly coming from data

Research examining AI pretraining progress from 2019 to 2025 indicates that data improvements have yielded 3.24 times more compute efficiency gains than model improvements. While model advancements were crucial for enabling larger-scale training, the primary driver of efficiency gains has been enhancements in data engineering. The gains from data and model improvements are largely independent, with data contributing 88% of the variance in performance.

  • From 2019 to 2025, AI pretraining saw 3.24x more compute efficiency gains from data improvements compared to model improvements.
  • Model improvements were crucial for enabling larger-scale training by overcoming constraints, rather than solely for compute efficiency.
  • Gains from data and model improvements are mostly independent, with additive effects explaining 88% of performance variance.
  • Data improvements, such as better filtering and larger scrapes, were key, while model progress included architectural and optimizer tweaks.
  • The effectiveness of data improvements may decrease for very large models with high capacity.
    Continue reading https://www.dwarkesh.com/p/pretraining-progress-is-mostly-data
Write a comment