CoreWeave SUNK

The industry's first unified training system for the most demanding AI workloads—delivering production-grade reliability and operational visibility for large, long-running training jobs.

Play video

Redefining the AI research cluster for production-grade training

SUNK is built for AI research teams running large, long-running training jobs, where predictability, reliability, and operational visibility matter as much as raw performance. SUNK preserves the Slurm workflows researchers rely on while bringing Kubernetes-native operational discipline to the cluster.

Lifecycle unity

Unify how researchers run Slurm and how platform teams operate clusters—without requiring weeks of bespoke setups. SUNK User Provisioning (SUP) automates secure onboarding and reduces identity/config drift so teams stay aligned from day one.

Reliability

Run large, long-running training jobs with production-grade reliability. CoreWeave Mission Control monitors cluster health end-to-end, detects silent hardware issues and GPU stragglers, and mitigates disruption before it compounds into lost training time.

Performance

Maximize productive training time with topology-aware scheduling and predictable cluster behavior tuned for distributed training. Keep multi-day runs moving forward by reducing disruption, retries, and fragmentation across GPU resources.

Observability

Get operational visibility from infrastructure health to job-level behavior. Correlate Slurm metrics with GPU, network, and storage signals to spot bottlenecks fast, validate performance, and keep training on track.

Proven by leading pioneers at production scale

A faster path to production-ready SUNK clusters

SUNK self-service streamlines how teams deploy and manage SUNK clusters capturing CoreWeave’s operational learnings from supporting research clusters of all sizes. Reduce setup friction, simplify operations, and move from cluster bring-up to productive training faster.

Play video

Run on industry-leading Cloud infrastructure services

SUNK runs on CoreWeave infrastructure services built for AI training performance, scale, and operational consistency.

Compute services

Get the latest GPU compute you need for your most complex AI workloads through a Kubernetes-native environment.

Storage services

Flexible, purpose-built, high-performance storage solutions that are purpose-built for AI.

Networking services

High-performance networking designed for optimal cluster scale-out and connectivity.

Supercomputing scale and enterprise-grade security

With massive megaclusters, CoreWeave GPU clusters help support multi-trillion parameter model training.

Left
Right

See what SUNK can do for you

Experience the resource flexibility your teams need to build, train, and deploy new models.

Zoho logo
CoreWeave
CoreWeave
Mistral AI logo
CoreWeave
Jane Street logo
Decart logo
Rev.com logo
Cloudflare logo
Abridge logo
CoreWeave
Fireworks AI logo
Augment logo
Conjecture logo
Chai logo
NovelAI logo
Runway logo

The market's most proven Slurm‑on‑Kubernetes offering

SUNK is the industry’s first unified training system for the most demanding AI workloads. SUNK is built for large, long-running training jobs where reliability and operational visibility matter.

SUNK self-service

Bring production-ready SUNK clusters online faster. SUNK self-service streamlines how teams deploy and manage clusters while reducing setup friction before productive training begins.

SUNK Anywhere

Extend the same unified training system beyond CoreWeave so teams can preserve one way of running demanding AI workloads as infrastructure environments expand. SUNK Anywhere helps reduce fragmentation across environments.

Mission Control

Monitor cluster health end to end with disciplined lifecycle control. Mission Control detects hardware anomalies and GPU stragglers and automatically mitigates failures to keep long-running training jobs on track.

Unified scheduling and observability

SUNK integrates Slurm and Kubernetes with tighter synchronization, unified scheduling, and built-in observability hooks. Reduce fragmented tooling and manual coordination as workloads scale.

Frequently asked questions

What is CoreWeave SUNK?

How does SUNK self-service help teams bring clusters into operation faster?

What is SUNK Anywhere?

How does Mission Control support reliability and operational visibility for long-running training jobs?

What performance and reliability outcomes has SUNK demonstrated?