Sized for pre-training, running post-training
At the frontier, the compute split has moved from 90 percent pre-training to roughly half. The gap shows up as queued rollouts, delayed evals, and iterations waiting on the loop’s slowest half.
Bursty fleets break both ways to buy
Sandbox and RL fleets spike to hundreds of thousands of concurrent environments per training step, then fall to near idle. Reserve for the peak and it sits empty; run on-demand and the training step waits.
Every workload is now a silicon bet
Published benchmarks rarely reflect how your own workloads behave. A wrong match is costly and sticky, paid back in porting work, stranded reservations, and re-validation.