CPU Compute

Post-training now claims roughly half of frontier compute, and nearly all of that new work runs on CPUs. CoreWeave runs it on bare metal: AMD EPYC and Intel Xeon today, NVIDIA Vera coming soon.

Your fleet was sized for pre-training. It's running post-training.

Models stopped just generating outputs and started doing work: running code, using tools, checking their own answers. That work runs on CPUs, at a scale nobody’s fleet was sized for.

Sized for pre-training, running post-training

At the frontier, the compute split has moved from 90 percent pre-training to roughly half. The gap shows up as queued rollouts, delayed evals, and iterations waiting on the loop’s slowest half.

Bursty fleets break both ways to buy

Sandbox and RL fleets spike to hundreds of thousands of concurrent environments per training step, then fall to near idle. Reserve for the peak and it sits empty; run on-demand and the training step waits.

Every workload is now a silicon bet

Published benchmarks rarely reflect how your own workloads behave. A wrong match is costly and sticky, paid back in porting work, stranded reservations, and re-validation.

Pack in more agents

How many environments you can run at once sets the ceiling on agentic AI. CoreWeave raises it: more agents per node, and per rack, with per-sandbox performance that holds.

No virtualization tax

No hypervisor overhead, no noisy neighbors, no shared tenancy. The throughput you validated in the PoC is the throughput your fleet delivers.

Built for the sandbox, down to the thread

Vera’s Spatial Multithreading gives every thread its own share of the processor, and a DPU on each node keeps infrastructure work off your sandboxes’ cores.

Silicon built for agent loops

Agent work is branch-heavy and memory-bound. Up to 1.8x faster agentic sandbox execution versus comparable x86, with 3x the memory bandwidth per core.

11,264
‍
cores per rack-scale configuration
22,528
‍
vCPUs per rack
11,000+
‍
concurrent environments
1.8
‍TB/s
NVLink network bandwidth per node
128 NVIDIA Vera CPUs in rack-scale configuration. Coming soon on CoreWeave.

Absorb every burst

A training step can need a hundred thousand environments for an hour, then almost nothing until the next run. Committed capacity carries the base, serverless catches the spikes.

Reserved base, serverless peaks

Steady load on committed CKS and SUNK capacity, spikes on serverless, through the same control plane as your GPU fleet. You buy the shape of your workload, not its worst hour.

Sandboxes inside the training platform

CoreWeave Sandboxes runs RL, agent tool use, and evaluation environments on the same infrastructure as your training. One operational surface at any burst size.

Spot economics for interruptible work

Rollout generation and preprocessing tolerate interruption. Spot at up to 60 percent savings, with a 7-minute preemption warning to checkpoint and drain cleanly.

10,000+
‍
nodes per production CKS cluster
Non-blocking
‍
InfiniBand and RoCE fabric
One network
‍
base and spillover scale without contention
Size the base for what you know and let the ceiling stay out of sight.

Every workload, optimized

Agent execution, RL environments, and data pipelines each run best on different silicon, and the right answer shifts as your loop evolves. Match each one without replatforming.

One platform, every answer

EPYC and Xeon bare metal today, Vera coming soon, all under the same CKS, SUNK, and Mission Control operating model. The workload moves between silicon; your tooling doesn’t.

ARM that matches your GPU estate

Vera brings ARM instances consistent with the hosts inside modern GPU systems, alongside CoreWeave’s x86 fleet. One toolchain, and the dual-validation tax goes away.

Choose it on guidance, not a bakeoff

Workload-fit guidance drawn from running agent, RL, and pipeline workloads at scale. Guided going in, reversible after: a working choice rather than a leap of faith.

CPU
Status
Architecture
Memory
Built for
CPU
NVIDIA Vera
Status
Coming soon
Architecture
ARM · 88 Olympus cores · Spatial Multithreading
Memory
Up to 1.8 TB/s NVLink
Built for
Agent sandboxes, RL rollouts
CPU
AMD EPYC™ 9005
Status
Available
Architecture
x86 · high-performance, high-core, general-purpose
Memory
Up to 1.5 TB per system
Built for
Up to 86% better gen-to-gen vs. EPYC 9004
CPU
Intel® Xeon® 5th Gen
Status
Available
Architecture
x86 · Platinum processors
Memory
512 GB per system
Built for
Up to 28% better than CoreWeave Xeon 3rd Gen
NVIDIA BlueField DPUs across the fleet. Footnotes carry each performance comparison.

Evaluate silicon on infrastructure someone else already validated

SemiAnalysis ClusterMAX™ Platinum

The only cloud rated Platinum, twice running.

SemiAnalysis on SUNK and CKS

Top rating for managing distributed AI workloads.

Constellation Research on CoreWeave Sandboxes

Purpose-built execution environments inside existing training infrastructure reduce operational sprawl and fragility.

The same operating model as your GPU fleet

CPU capacity runs under CKS, SUNK, CoreWeave Sandboxes, and CoreWeave Mission Control, the control plane your GPU clusters already use.

CKS

Managed Kubernetes for AI.

SUNK

Slurm on Kubernetes.

Sandboxes

Isolated environments for RL and agent tool use.

Mission Control

Fleet health and observability.

The CPU work in your AI loop, with GPU-grade discipline

Sandboxes, RL environments, and data pipelines on bare metal. Committed capacity for the base, serverless for the spikes, spot for interruptible work. NVIDIA Vera coming soon.