Under the Hood: What Predictable Inference Actually RequiresUnder the Hood: What Predictable Inference Actually RequiresUnder the Hood: What Predictable Inference Actually Requires
CoreWeave

Under the Hood: What Predictable Inference Actually Requires

Inside Inference Live Webinar

Event details

Location
Sitanshu Gupta
Director of Engineering, Inference Services, CoreWeave
,
CoreWeave
Location
Zachary Xi
Founding Product Manager
,
Inferact
Location
Schedule

Sep 17, 2026

11:00 am

ET

September

17

 — 

Location
45 minutes

What changes when a model actually has to handle production load?

Most inference demos run at low concurrency, where latency is stable and cost is an afterthought. Production traffic breaks that picture fast: request bursts, cache pressure, and throughput/latency tradeoffs that only show up at scale.

In this session, vLLM talks through the scheduling, memory management, and quantization decisions behind their serving stack. CoreWeave covers the infrastructure layer—networking, topology, and orchestration—that determines whether improved throughput and latency show up as real performance gains or get bottlenecked before they reach the model.

In this webinar, we'll cover:

  • Why inference demos break at production scale—and how request bursts, cache pressure, and throughput/latency tradeoffs create behavior that stays hidden in stable, low-concurrency tests.
  • How vLLM serving works in production and the vLLM roadmap
  • How CoreWeave infrastructure—including GB300 NVL72, networking topology, and orchestration—supports vLLM performance
  • What must be true across the hardware and platform layers for scheduling and batching gains to show up as measured performance rather than being absorbed by hardware limits.

Speakers

Sitanshu Gupta
Sitanshu Gupta
CoreWeave
Director of Engineering, Inference Services, CoreWeave
Zachary Xi
Zachary Xi
Inferact
Founding Product Manager

CoreWeave Cloud,
Home v3,
Home v2,
Product - GPU Compute,
Product - Virtual Servers,
Solution - Pixel Streaming,
Solution - Machine Learning,
Product - VFX,
Product - Kubernetes,
Product - Concierge Render,
Home,