RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor
Sequoia Capital·2026-08-12 13:00

Financial Performance & Market Trends - Revenue run rate grew from $1 billion to $2 billion over the course of approximately 4 months[1] - Custom task pricing ranges from $50 to $10,000, with certain frontier labs purchasing 50,000 tasks per month[25][30] - Expert network throughput reached 2.5 million hours in Q2 alone[10] - Post-training compute investments for specific task datasets reached approximately $500k (thousand)[20][21] Data Evolution & Technical Strategy - Industry transition shifted from the low-skilled crowdsourcing era of behavior cloning to the agentic era of data utilizing high-skilled experts like software engineers, lawyers, doctors, and bankers[4][5] - Reinforcement Learning (RL) environments are structured around three core pillars: worlds (documents and sheets), apps (high-fidelity software clones like Salesforce and Microsoft 365), and tasks (prompts and verifiers)[8] - Corporate law benchmark performance increased significantly from 4.7% to 26.6% after targeted post-training utilizing expert-curated agent data sets[21] - Data curation and pricing strategies are categorized into three main models: bespoke task pricing, off-the-shelf data investments, and hourly expert models[26][27]

RL Environments Explained: How AI Agents Learn Real-World Work | Brendan Foody, Mercor - Reportify