Serverless
Provider 
Founded 2016
Sells To Developer
Pricing Model Recurring Usage-Based
Ownership Privately Held

Vast.ai Serverless runs autoscaling GPU inference endpoints on managed capacity.
Vast.ai Serverless deploys models as autoscaling inference endpoints, with optimization that selects hardware across the Vast.ai fleet and scales to zero between requests. Deployments are defined in Python through the Vast SDK.
Buyers pay only for compute time while handling inference traffic, without managing GPU capacity, tiers, or minimum commitments.
| Attribute | Serverless | Dedicated Inference | Modal Cloud | Replicate Cloud | |
|---|---|---|---|---|---|
| Provider | |||||
| Founded | 2016 | 2019 | 2021 | 2019 | |
| Sells To | Developer | Enterprise | Enterprise | Enterprise, Small Business | |
| Pricing Model | Recurring Usage-Based | Recurring, Software, Usage-based | Recurring, Software | Recurring, Software, Usage-based | |
| Ownership | Privately Held | Equity, Venture Capital | Venture Capital | Venture Capital |
Provider 
Founded 2016
Sells To Developer
Pricing Model Recurring Usage-Based
Ownership Privately Held
Provider 
Founded 2019
Sells To Enterprise
Pricing Model Recurring, Software, Usage-based
Ownership Equity, Venture Capital
Provider 
Founded 2021
Sells To Enterprise
Pricing Model Recurring, Software
Ownership Venture Capital
Provider 
Founded 2019
Sells To Enterprise, Small Business
Pricing Model Recurring, Software, Usage-based
Ownership Venture Capital

Hugging Face Inference Endpoints

Cloudflare Workers
cloudflare.com

Together AI Inference
together.ai

Edge AI Inference
cloudflare.com

Baseten Inference Stack

Cerebras Cloud (Inference & Training Platform)
cerebras.ai

Spheron GPU Cloud Marketplace
spheron.network

Inference API

Serverless Container Platform

Aptible LLM Gateway

RunPod Cloud

Cloudflare AI Gateway