
Post-Training pricing
Your loop. Your stack. Your call.
Serverless SFT and RL
Training is priced per GPU-hour and prorated by active training time. Each GPU comes pre-provisioned with the base model weights and training image, so you pay only for active training time. Inference, evaluation, and checkpoint storage are billed separately.
Model Distillation
Free during Private Preview
Model Distillation uses Serverless SFT to train a smaller, faster model to reproduce a larger model’s behavior on a specific task. See the documentation to learn when to use Model Distillation.
Model Distillation is not billed separately, but uses hosted CoreWeave services to collect traces, generate datasets, train models, and serve them for inference.¹
Frequently asked questions
How much does a typical training job cost?
Costs vary by model, dataset, and training settings. At $2.70 per hour, 20 minutes of billable training costs $0.90, one hour costs $2.70, and four hours costs $10.80. These are illustrative training costs; inference, evaluation, and storage are additional.
What factors affect the cost of an RL or SFT training job?
Training costs depend on how long the job runs, at $2.70 per hour. Model size, sequence length, and the amount of training data affect runtime. For SFT, additional epochs increase training time. For RL, generating more or longer rollouts increases inference costs and can increase training time. Reward-model calls, evaluation, and storage can add separate costs.
How is training usage calculated?
Serverless SFT and RL both cost $2.70 per hour of active training, prorated by usage. For example, 30 minutes of billable training costs $1.35. You aren’t charged for idle time. Inference and checkpoint storage are billed separately.
How does pricing for Model Distillation work?
Model Distillation is free during its preview period. You will not be billed for its usage during preview. We’ll communicate pricing before billing begins.
Where can I view my monthly token usage?
Your CoreWeave Forge billing dashboard shows inference token usage for the current and previous months. During preview, Model Distillation usage is shown for visibility only and is not billed.
What services contribute to Model Distillation usage?
Model Distillation has no separate fee. It may use several CoreWeave Forge services, including teacher-model inference, dataset generation, student training through Serverless SFT, evaluation, and artifact storage. During Private Preview, you will not be charged for usage of these services when used as part of the Model Distillation workflow.