
Aranya converts bare-metal GPU servers into managed, production-ready AI inference clusters within 48 hours.
clusterdOS is the open-source engine that runs inside each cluster, licensed GPLv3-or-later and installable from a public GitLab repository. It composes Kubernetes with Cilium networking, Ceph and WekaFS storage, Grafana and Prometheus observability, NVIDIA and AMD GPU operators, and vLLM-based inference serving, all configured declaratively through ArgoCD.
Above that engine, the Aranya multicluster operating system federates clusters so a whole fleet can be reasoned about and directed as one system, including in plain language. A fully managed tier adds architecture design, deployment, capacity scaling, and around-the-clock incident response so customers need not staff a platform team.
Aranya argues that inference has become the dominant AI workload, citing a projection that it reaches two-thirds of all compute during 2026 from roughly a third three years earlier. It frames the gap between raw data-center capacity and production-ready AI infrastructure as the constraint that this growth exposes.
Its supply-side target is data centers and neoclouds at 20 megawatts and below, which it says hold hardware but lack the platform layer needed to sell that hardware as AI compute. On the demand side it targets inference and training teams choosing between hyperscaler pricing and building a cloud from scratch on bare metal.
The stated differentiator is operating below the workload layer: rather than only rescheduling a pod off a failing node, clusterdOS detects and resolves GPU thermal events, ECC errors, and networking faults on the hardware itself. That hardware-level scope is what the company credits for cutting one customer's cluster setup from six weeks to under 48 hours and reducing outages by 90%.
The free open-source core lowers the cost of trial and builds a community around the same engine the paid service runs, while the managed tier competes on price against hyperscaler Kubernetes offerings. Named production deployments with Baseten and Hydra Host give a company barely a year old reference customers in demanding inference environments.
Aranya publishes three tiers. clusterdOS is free to self-host as open source with community Slack support, the Aranya multicluster operating system is not yet released and carries no announced price, and fully managed clusters are quoted at $0.06 to $0.16 per GPU-hour depending on requirements.
Charging by GPU-hour ties revenue to a customer's fleet size rather than to seat count, and managed storage is sold as a separate add-on. The free engine functions as the entry point, with the managed tier carrying the support, on-call rotation, and observability dashboards that justify the price.