INFERENCE WITHOUT THE INFRASTRUCTURE WORK

Serverless Inference

Call an API and start building. No clusters, no GPU selection, no infrastructure to stand up.

Why Serverless Inference

Start building without standing up infrastructure

With Serverless Inference, you call an API and get instant access to leading open-source models and your own fine-tuned checkpoints. CoreWeave hosts the models and runs the GPUs, so the only thing you manage is the application.

Leading open-source models or your own LoRA weights, one API

Access a curated catalog of open-source models, or bring your own LoRA weights and serve them side by side. Switch between them without changing your integration.

Tracing, evals, and observability

Every request can be traced and evaluated through native observability features, so teams can debug and improve AI applications and agents without extra instrumentation.

Playground access with zero configuration

Explore and compare models directly in the playground before writing a line of code. No endpoints to configure and no keys to manage.

Models available on serverless inference

DeepSeek V4.1-Flash

New
Sep 2026
$0.20 input / $0.03 cached / $0.65 output
1M

DeepSeek V4.1 Flash is a multimodal Mixture-of-Experts model supporting text and image inputs. It supports coding, reasoning, and agentic tasks with long context.

IBM Granite 4.2 8B

New
Aug 2026
$0.10 input / $0.05 cached / $0.15 output
131K

Granite 4.2 8B is a performant instruct model for its size capable of enhanced tool calling, instruction following, and chat capabilities.

Qwen3.8 27B

New
Aug 2026
$0.40 input / $0.15 cached / $3.00 output
262k

Qwen3.8-27B is a dense multimodal model suited for coding, research, vision, and long-running agent tasks with broad support for popular harnesses and development tools.

‍

Z.AI GLM 5.3 Flash

New
Aug 2026
$0.15 input / $0.05 cached / $0.50 output
1M

GLM-5.3-Flash is a natively multimodal model with 320B total parameters and 18B active parameters excelling in agentic tasks, coding, and tool use.

DeepSeek V4-Pro-0813

New
Aug 2026
$1.31 input / $0.044 cached / $3.96 output
1M

DeepSeek V4-Pro-0813 is a 1.6T-parameter MoE model excelling at advanced reasoning, coding, and complex agentic workloads.

NVIDIA Nemotron 3.5 Lightning

New
Aug 2026
$0.07 input / $0.04 cached / $0.20 output
262K

Nemotron 3.5 Lightning is an MoE model built for fast, reliable agentic tasks across use cases such as financial service

Who Serverless Inference is built for

Move fast without standing up infrastructure

The teams that get the most value from Serverless Inference are testing, iterating, and shipping AI applications faster than they can provision infrastructure for them.
AI engineers

Iterating on AI applications and agents

Teams building copilots, agents, and RAG applications who need to test and swap models quickly without provisioning infrastructure for every experiment.
AI teams

Bringing fine-tuned models to production fast

Teams that have fine-tuned a model using LoRA and need to serve it immediately, without building a dedicated deployment pipeline first.
Platform teams

Evaluating a workload before committing to scale

ML platform teams who need to grow into distributed inference while preserving GPU choice, runtime flexibility, and per-GPU-hour cost visibility as workloads expand.
CoreWeave inference paths

Inference on your terms

Three inference paths built on the award-winning CoreWeave Cloud. Move between them as your workloads evolve—without replatforming—so you consistently get predictable performance and infrastructure-aligned economics.
Serverless Inference
Pay-per-token inference on a curated OSS model catalog. No clusters to manage. Built-in tracing, evals, and observability for AI applications and agents.
Best for
Rapid iteration and AI app development
Infrastructure management
Low — API only
Model support
Curated OSS + LoRAs
Pricing
Pay-per-token
Dedicated Inference
Deploy custom weights on explicitly chosen GPUs. CoreWeave operates routing, scaling, and lifecycle. Full execution transparency, no cluster management.
Best for
Custom model serving at production scale
Infrastructure management
Medium — GPU, Zone, Runtime
Model support
Open-source or custom weights
Pricing
Pay-per-GPU-hour
Inference on CKS
Full self-managed inference on CoreWeave Kubernetes Service. Own the entire serving stack—runtimes, scheduling, autoscaling, multi-node topology—on dedicated bare-metal GPU nodes.
Best for
Full infrastructure ownership and deep tuning
Infrastructure management
High — full Kubernetes control
Model support
Any
Pricing
Pay-per-GPU-hour (Reserved, On-demand, Spot, Flex)
INFERENCE IS PART OF COREWEAVE FORGE

Serve models where your training runs

CoreWeave Forge connects the AI loop and keeps you free to build with any cloud, model, or framework, so improvement compounds with every version. Traces from training, scores from evaluation, and your live inference all run on Forge, with no tool handoffs and no rebuilding experiment context by hand.

Run

Improve

Evaluate

HOW IT HELPS

What can you do with Serverless Inference?

Go from open-source model to live endpoint in minutes

Call the API or use the playground. No signing up with a separate hosting provider, no cluster to provision, no GPU to select. Point your application at the endpoint and start sending requests.

Iterate on fine-tuned models without managing serving infrastructure

Bring your own LoRA weights and serve fine-tuned models without building a pipeline to fetch, hot-swap, and scale weights for every iteration. CoreWeave handles that complexity so teams can move between training and inference quickly.

See how your application behaves in production, not just at the prompt

Catch the request that failed in production, not just the one that failed at the prompt. See exactly where an agent’s reasoning broke down, days after it shipped.

How Serverless Inference works

Four steps to your first API call

Choose a model, call the API, and start iterating. There is no cluster to provision and no GPU to select.

1. Choose a model

Pick from the curated catalog of open-source models, or point to your own LoRA weights.

2. Call the API

Send requests to the OpenAI-compatible endpoint using your existing SDK. No new client library to learn.

3. Trace and evaluate

Every request can be traced automatically. Review, evaluate, monitor, and iterate on AI applications and agents directly, no instrumentation or setup required.

4. Scale as usage grows

Pay per token as traffic increases, monitor your spend in the billing dashboard and set an inference budget. No capacity to forecast, no headroom to buy in advance.

Related resources

Blog

GLM 5.2 Available on Serverless Inference

Blog

Production AI Runs on Inference. Are You Ready for It?

Solution brief

Choosing the Right Execution Path for AI Inference

Frequently asked questions

Which models are available through Serverless Inference?

Can I run my own fine-tuned model on Serverless Inference?

How is Serverless Inference billed?

Are there usage limits or concurrency caps I should plan around?

How does tracing and observability work?

Can I switch between models without redeploying anything?

Can I evaluate a model before integrating it into my application?