
Modal Labs runs a serverless cloud platform for AI inference, training, sandboxes, and batch workloads.
Modal Labs operates a serverless cloud platform designed for AI, machine-learning, and data-intensive workloads. Customers write Python code that defines functions, containers, and hardware requirements; Modal handles provisioning, scheduling, scaling, and observability across global GPU and CPU capacity.
The platform supports inference, training, batch processing, and secure sandboxes. Modal routes workloads across clouds and regions in real time, scales from zero to thousands of GPUs on demand, and provides integrated logging, metrics, and security controls including SOC 2 and HIPAA compliance.
Modal differentiates itself with sub-second container cold starts and instant autoscaling, which reduces wait times and cost for bursty AI inference and training jobs. Its AI-native runtime is purpose-built for GPU-heavy workloads across a globally distributed fleet, removing the need for customers to reserve capacity or manage complex orchestration.
The single-codefile experience lets teams specify logic, dependencies, and hardware in one Python file, lowering infrastructure complexity. Integrated observability, multi-cloud routing, and pay-by-second billing reinforce the platform's positioning as a developer-centric alternative to self-managed Kubernetes or hyperscaler compute services.
Modal uses a pay-by-second consumption model with no long-term capacity commitments, so customers pay only for the compute they use. Workloads scale to zero between requests, and the platform routes jobs across GPU types and regions to match demand.
Modal offers a $30-per-month free compute tier and publishes per-second GPU/CPU pricing. This structure targets cost-sensitive AI startups and research teams while supporting enterprise workloads that need fine-grained cost visibility across inference, training, and sandbox workloads.