REINFORCEMENT LEARNING

Serverless RL

Post-train LLMs for multi-turn agentic tasks while keeping control of examples, environments, rewards, and hyperparameters. CoreWeave manages elastic GPU capacity and distributed training so you can focus on agent behavior.

Reinforcement learning for agents that learn from outcomes

Not every agent problem needs reinforcement learning (RL). Prompts, tools, and context are faster to change and often resolve the issue. RL is for what instruction can’t fix: multi-turn tasks where success depends on a sequence of decisions, and the only way to improve the agent is to let it try and score how it did.

That takes three things: a way to run the agent and capture its performance, a way to judge which attempts went better, and somewhere to train.

Serverless RL

Run reinforcement learning workloads on managed CoreWeave infrastructure that scales with training, then back to zero when the work is complete.

Agent Reinforcement Trainer

Build, evaluate, and iterate on multi-turn agents with an open-source framework that connects your application code to reinforcement learning algorithms and Serverless RL.

RULER

Generate relative rewards with an LLM-as-judge approach that ranks groups of agent trajectories, without labeled data or a handcrafted reward function.

SERVERLESS RL IS PART OF COREWEAVE FORGE

Turn scored outcomes into better agent behavior

CoreWeave Forge connects the traces, evaluations, training, and inference behind your AI application. Serverless RL uses scored trajectories to improve the decisions an agent makes across multi-turn tasks, then returns the trained model to inference for testing and continued iteration.

Improve

Frequently asked questions

What is Serverless RL?

When should I use Serverless RL?

Do I need to provision a training cluster?

What is Agent Reinforcement Trainer?

What is RULER?

Can I train from production traces?

How does Serverless RL improve training efficiency?

Can I move between Serverless SFT and Serverless RL?

Related resources

The same news reads differently depending on where you sit. Here’s the version that applies to you.

BLOG

CoreWeave Forge: Turn AI Iteration Into Compounding Improvement

VIDEO

CoreWeave Forge Explainer Video

PRESS RELEASE

CoreWeave Forge Press Release

Builder Resource Center: Learn from every run

Explore demos, code, and technical resources for every stage of the AI loop. Learn how researchers, developers, and CoreWeave engineers build, observe, evaluate, and improve AI systems—and put those insights to work.

GET STARTED

Let your agents learn from every run

Bring your agent, environment, and reward. Serverless RL handles the GPUs, distributed training, and scaling, and RULER scores trajectories without a handcrafted reward function. Copy the notebook, run your first training job, and send the improved checkpoint back to inference.