Industry Trends & Challenges - AI progress faces growing time-consuming and expensive benchmark scaling, taking anywhere from 4 hours to several days [3] - Current algorithms suffer from a task distribution mismatch, off-policy sampling, massive parallel rollout infrastructure bias, and single-scalar reward bottlenecks [6][7][8] - The industry handles a massive token problem, spending hundreds of trillions of tokens daily on inference and generating real-world failure and success data [4] Investment Opportunities & Technological Solutions - Trajectory developed a continual learning platform focusing on turning production agent traces into model and harness improvements [48][49] - On-Policy Self-Distillation (OPSD) solves reinforcement learning limitations by achieving online task distribution, on-policy sampling, single parallelism, and per-token dense rewards [23][24] - OPSD successfully scales up to 120 billion (120B) parameter models, 500 billion (500B) parameter models, and 1 trillion (1T) parameter models for complex tool-calling tasks [29] - Advanced algorithmic solutions like step-level divergence weighting and residual guidance are implemented to solve long-horizon task degradation and hint leakage issues [34][42][43] Company Financials & Performance - The foundational model sui 1 led to a 2 billion (2B) dollar acquisition at DeepMind [2] - Trajectory successfully scales its self-distillation algorithm up to 12 billion (12B) parameter models on Mercury Apex agents requiring 100 or more tool calls [46] - Early access to the platform has been deployed to key industry companies including Harvey, Dakagon, and Rogo [51]
Scaling up Continual Learning — Ronak Malde, Trajectory
AI Engineer·2026-08-12 14:30