Pricing

Post-Training pricing

Your loop. Your stack. Your call.

Serverless SFT and RL

Training is priced per GPU-hour and prorated by active training time. Each GPU comes pre-provisioned with the base model weights and training image, so you pay only for active training time. Inference, evaluation, and checkpoint storage are billed separately.

Training method
Price ($/GPU hr)
Context limit
Inference usage
Training method
SFT
Price ($/GPU hr)
$2.70/hr
Context limit
32K
Inference usage
Optional dataset generation and evaluation, billed separately
Training method
RL
Price ($/GPU hr)
$2.70/hr
Context limit
32K
‍Inference usage
Rollout generation and evaluation, billed separately

Model Distillation

Free during Private Preview

Model Distillation uses Serverless SFT to train a smaller, faster model to reproduce a larger model’s behavior on a specific task. See the documentation to learn when to use Model Distillation.

Model Distillation is not billed separately, but uses hosted CoreWeave services to collect traces, generate datasets, train models, and serve them for inference.¹

FAQS

Frequently asked questions

How much does a typical training job cost?

What factors affect the cost of an RL or SFT training job?

How is training usage calculated?

How does pricing for Model Distillation work?

Where can I view my monthly token usage?

What services contribute to Model Distillation usage?

¹ Services used by Model Distillation may include CoreWeave Agent Lens, Weights & Biases Models, Serverless SFT, Serverless Inference, and artifact storage.