Ollama was founded in 2021 by Jeffrey Morgan and Michael Chiang, who previously co-founded Kitematic, an early graphical interface for Docker that Docker Inc. acquired in 2015. The company participated in Y Combinator's Winter 2021 batch after an initial $125,000 pre-seed.
Ollama later raised a $15 million Series A led by Benchmark's Peter Fenton, who joined the board, and in July 2026 closed a $65 million Series B led by Theory Ventures with participation from Benchmark, 8VC, Y Combinator, Pace Capital, Garage Capital, GTMFund, and 49 Palms, bringing total funding to $88 million. By mid-2026 the company had about 14 employees and 8.9 million monthly active developers.
Ollama provides a command-line interface and REST API to download, configure, and run open-source models while handling quantization, acceleration, and dependencies. It supports CPU-only setups and GPU acceleration across Apple Silicon, NVIDIA, and AMD.
Ollama includes official Python and JavaScript libraries, Docker support, and integrations with major coding tools including Claude Code, OpenCode, and Codex. The platform also offers cloud-hosted inference with Free, Pro ($20/mo), and Max ($100/mo) tiers.
Ollama sits at the intersection of several major trends: privacy-preserving AI, open-source model proliferation, and developer tool consolidation. With 172,000+ GitHub stars and recognition as the fastest-growing open-source startup of 2024, Ollama has established itself as the default local inference platform.
Strategic integrations include Google Firebase Genkit and Gemma models, OpenAIs GPT-OSS family, Anthropic Claude Code compatibility, and Apple MLX framework on Apple Silicon. The company is expanding beyond core infrastructure into consumer applications with OpenClaw and experimental image generation. The local LLM market is growing as enterprises seek data sovereignty and developers demand alternatives to cloud API costs.
Ollama is the largest developer network in the open model ecosystem, with 8.9 million monthly active developers and more than 67,000 community-built integrations on GitHub. It is used within 85% of the Fortune 500, including customers in regulated industries such as government, healthcare, and finance.
Ollama is a distribution and early-access partner for the major open model labs, including Meta, Google DeepMind, Mistral, and MiniMax, and for hardware vendors including NVIDIA, Intel, AMD, and Qualcomm, giving users day-zero access to new models. It offers a local-first, privacy-preserving experience with an OpenAI-compatible API and no training on user data.
Ollama does not provide a native graphical user interface, making it less accessible to non-technical users compared to LM Studio or GPT4All. Local deployment requires adequate GPU hardware and DevOps knowledge for setup and maintenance.
The cloud model catalog is smaller than dedicated API providers like Together or Fireworks. Cloud subscription tiers operate on session limits that reset every 5 hours and weekly cycles, which can be restrictive for bursty or sustained high-volume workloads. Per-token pay-as-you-go billing is listed as coming soon but not yet available.
Ollamas pricing follows a dual-track freemium model. The local open-source software remains free with no license cost, unlimited public models, and no usage limits on localhost.
Ollama Cloud offers three subscription tiers: Free (1 concurrent cloud model, basic limits), Pro ($20/month or $200/year, 3 concurrent models, 50x more cloud usage, private model uploads), and Max ($100/month, 10 concurrent models, 5x Pro usage). Cloud billing is based on GPU utilization time rather than per-token consumption, which means efficiency gains from newer hardware directly benefit users.

Ollama is an open-source platform for running open-weight AI models locally or in its cloud.