Pre-training Data Composition and Paradigm Shifts - Traditional base models relied heavily on general human knowledge, with WebText and Wikipedia comprising roughly 85% of the training mix in GPT-3 [5] - Recent models like Llama 3 allocate about 50% of their training tokens to general knowledge [5] - The proportion of raw web text has significantly decreased, dropping down to 15% in newer model architectures [12] - Early pre-training recipes lacked specific code data sets, whereas code has now become a dominating data subset [16] - Industry practices now incorporate supervised fine-tuning (SFT) and question-answer chat data early into the pre-training phase rather than keeping it strictly for post-training [14] Reinforcement Learning and Compute Allocation - Industry adoption of reinforcement learning (RL) has shifted from being a minor enhancement to dramatically improving model performance on diverse tasks, pioneered by models such as OpenAI o1 in 2024 and DeepSeek R1 in January 2025 [8] - Xiaomi's Mimo Labs allocates roughly equal compute budgets between pre-training and post-training in the final model, while some architectures sink significantly more compute into RL than supervised learning [23] - Base models are increasingly viewed as providers of atomic skills and reasoning priors to build effective foundations for large-scale reinforcement learning and agentic behavior rather than acting as static web text repositories [24][28]
The Base Model Is Dead — Varun Singh, Arcee AI
AI Engineer·2026-07-31 20:30