Serverless Inference
Call an API and start building. No clusters, no GPU selection, no infrastructure to stand up.
Start building without standing up infrastructure
Models available on serverless inference
Move fast without standing up infrastructure
Iterating on AI applications and agents
Bringing fine-tuned models to production fast
Evaluating a workload before committing to scale
Inference on your terms
Serve models where your training runs
CoreWeave Forge connects the AI loop and keeps you free to build with any cloud, model, or framework, so improvement compounds with every version. Traces from training, scores from evaluation, and your live inference all run on Forge, with no tool handoffs and no rebuilding experiment context by hand.

Run
Serverless and Dedicated Inference power the Run stage of the AI loop.
Choose Serverless to start fast, pay per token, and scale instantly. Pick Dedicated for high-volume workloads, reserve compute, and lock in capacity.
Improve
Evaluation results show what to fix, so you can feed them into your next training run. The new version deploys back to Serverless or Dedicated Inference.
Evaluate
Production requests generate traces that show what your model actually does. Forge scores those traces against your evaluations, so you measure quality on live traffic behavior.

What can you do with Serverless Inference?
Go from open-source model to live endpoint in minutes
Call the API or use the playground. No signing up with a separate hosting provider, no cluster to provision, no GPU to select. Point your application at the endpoint and start sending requests.
Iterate on fine-tuned models without managing serving infrastructure
Bring your own LoRA weights and serve fine-tuned models without building a pipeline to fetch, hot-swap, and scale weights for every iteration. CoreWeave handles that complexity so teams can move between training and inference quickly.
See how your application behaves in production, not just at the prompt
Catch the request that failed in production, not just the one that failed at the prompt. See exactly where an agent’s reasoning broke down, days after it shipped.

Four steps to your first API call
Choose a model, call the API, and start iterating. There is no cluster to provision and no GPU to select.
1. Choose a model
Pick from the curated catalog of open-source models, or point to your own LoRA weights.
2. Call the API
Send requests to the OpenAI-compatible endpoint using your existing SDK. No new client library to learn.
3. Trace and evaluate
Every request can be traced automatically. Review, evaluate, monitor, and iterate on AI applications and agents directly, no instrumentation or setup required.
4. Scale as usage grows
Pay per token as traffic increases, monitor your spend in the billing dashboard and set an inference budget. No capacity to forecast, no headroom to buy in advance.
Frequently asked questions
Which models are available through Serverless Inference?
Can I run my own fine-tuned model on Serverless Inference?
How is Serverless Inference billed?
Are there usage limits or concurrency caps I should plan around?
How does tracing and observability work?
Can I switch between models without redeploying anything?
Can I evaluate a model before integrating it into my application?
Inference built on the Essential Cloud for AI
Serverless Inference gives you instant access to leading open-source models, per-token economics, and observability built in. Get started today.


