As AI models evolve, inference throughput must be evaluated and validated in context. A token generated by a text-only model does not place the same demands on infrastructure as one generated by a multimodal or MoE reasoning model. Multimodal architectures add compute-intensive image encoding and cross-attention, while MoE reasoning models route tokens through specialized experts and often generate longer responses. These differences change the compute, memory, and networking required to sustain a given tokens-per-second rate and ultimately what that throughput costs.
With MLPerf® Inference v6.1, throughput is evaluated across a diverse set of model architectures. CoreWeave participated in the Datacenter Closed division’s Available category, submitting results across four NVIDIA platforms and four model families that are in use today. NVIDIA Blackwell and Blackwell Ultra are commonly adopted platforms in production today. The models benchmarked are the following:
- Qwen3-VL-235B-A22B, multimodal AI powering production search and catalog use cases
- DeepSeek-R1-671B, one of the largest MoE reasoning models and widely adopted
- GPT-OSS-120B, the open-weight reasoning model seeing rapid adoption since release
- Llama 2 70B, one of the most widely deployed models in production
Together, these throughput numbers on the infrastructure and models are already doing the work in production and delivering real world performance.
What the results mean for customers
Benchmarks capture performance at a moment in time, but the capabilities behind that performance continue to compound. CoreWeave extracts more value from deployed hardware while rapidly bringing the latest silicon and emerging model classes to run at production-scale.
For customers evaluating inference infrastructure, two measures matter most: tokens per GPU1 and time to production. Higher per-GPU throughput can reduce the infrastructure required to serve a workload, while rapid access to new platforms helps teams move from model development to production sooner.
CoreWeave’s performance across multimodal, frontier-scale reasoning, MoE, and large language models gives customers the efficiency and scale to turn NVIDIA accelerated computing into real world AI work.
CoreWeave delivers the highest Qwen3-VL-235B-A22B server throughput among cloud providers
CoreWeave’s results with Qwen3-VL-235B-A22B highlight how our software stack turns new multimodal scenarios and rack-scale infrastructure into leading inference performance. CoreWeave delivered the highest server throughput among cloud providers using an NVIDIA GB300 NVL72 rack. In server scenario, CoreWeave sustained 1,196 queries per second, the highest server throughput among cloud providers, and 1,135 samples per second in offline scenario.
CoreWeave increases DeepSeek-R1-671B per GPU throughput by nearly 20% over MLPerf® v6.0
Continuous software optimization increases the throughput of GPUs already in production, helping customers serve more tokens and improve inference economics without waiting for the next hardware generation. Significant gains can come from regular tuning that delivers more throughput which is why CoreWeave continuously optimizes its serving stack, scheduling, and operations to unlock more performance from existing infrastructure. The results are measurable. Just five months after MLPerf® Inference v6.0, comparing the 72-GPU v6.1 submission with the 64-GPU v6.0 submission on NVIDIA GB200 NVL72,derived per-GPU Server throughput increased by 19.8%.
CoreWeave achieves the highest GPT-OSS-120B per GPU throughput
On a single NVIDIA GB300 NVL72 rack, CoreWeave sustained 16,635 tokens per second per GPU in offline scenario, and 16,118 in server scenario. These were the highest per-GPU throughput of any v6.1 Datacenter Closed submission on GPT-OSS-120B, on any silicon, in both scenarios.1 CoreWeave achieved over 1.16 million tokens per second in server and over 1.19 million tokens per second in offline scenarios on a single NVIDIA GB300 NVL72 rack.
Per-GPU throughput maps directly to inference economics: At a given GPU price, serving more tokens per accelerator can reduce infrastructure cost per token. CoreWeave’s advantage also extends across platforms: Our NVIDIA GB200 NVL72 posted the highest rack-scale GPT-OSS-120B total in its class, reaching 901,058 tokens per second in server and 911,566 tokens per second in offline scenarios. Our NVIDIA HGX B200 and HGX B300 systems led all HGX submissions from cloud providers in both scenarios.
CoreWeave delivers the highest throughput among cloud providers with Llama 2 70B
With Llama 2 70B, CoreWeave demonstrated rack-scale serving capacity for a workhorse model that is widely deployed. We delivered the highest throughput among cloud providers with NVIDIA GB300 NVL72 with a throughput of 944,902 tokens per second in server scenario and 1,136,100 in offline scenario. In the offline scenario, we achieved the only result above one million tokens per second among cloud provider submissions using NVIDIA GB300 NVL72.
Summary of CoreWeave’s Performance Results
CoreWeave’s MLPerf® Inference v6.1 submissions covered four distinct model families with multiple NVIDIA Blackwell and Blackwell Ultra platforms:
Benchmarks were performed on production clusters
These results were not produced on a tuned cluster optimized just for benchmarking. CoreWeave ran its submission on the same production clusters, images, and operational stack available to customers:
- CoreWeave Mission Control provides fleet-wide health monitoring and lifecycle management, with continuous benchmarking that ties system behavior to sustained tokens-per-second outcomes.
- Topology-aware scheduling via SUNK pins inference workloads inside high-bandwidth NVIDIA NVLink domains on NVIDIA GB200 NVL72 and GB300 NVL72 systems.
The same discipline that lets CoreWeave stand up four Blackwell platforms—NVIDIA HGX B200, HGX B300, GB200 NVL72 and GB300 NVL72 in one submission window is what moves customers from delivery to serving tokens in days.
To learn more:
Join us at Fully Connected
Watch Many Workloads, One Cluster: Training and Inference Without Idle GPUs
Read the Signal65 report on What does AI infrastructure really cost?











