Volantis Inc. today announced that it has raised a $88 million funding round led by prominent angel investor Lachy Groom and Abstract Ventures.
The Series A deal also drew more than a half-dozen other participants. The group included Kleiner Perkins chair John Doerr and Naveen Rao, the former head of Intel’s artificial intelligence products group.
Volantis’ founding team likewise has strong industry credentials. The company employs engineers who previously worked at major chipmakers such as Nvidia Corp. and Broadcom Inc. Their technical achievements include the first commercial implementation of CoWoS, an interconnect widely used in graphics processing units.
The speed at which data travels between a GPU’s processing cores and HBM memory is a major contributor to AI model performance. It’s measured by a metric called memory bandwidth. The more memory bandwidth a GPU has, the faster it can perform inference.
Volantis is developing an inference-optimized chip architecture that is built around a custom memory interconnect. According to the company, the technology provides more than 30 times the memory bandwidth of current accelerators.
The simplest way to increase a GPU’s memory bandwidth is to add more memory modules. Today, the number of memory modules that can be placed on a GPU is constrained by the wires that link them to the GPU’s processing cores. Those wires are up to 5 millimeters long. Memory modules must be placed within the wires’ 5 millimeter range, which heavily limits the total number of modules that can fit on a chip.
Volantis’ chip architecture features interconnect wires with a range of over 200 millimeters. According to the company, that extended range makes it possible to equip an AI accelerator with more than 220 memory chiplets. The result is a significant increase in memory bandwidth.
Volantis’ memory interconnects are based on an optical design, which means that they transmit data in the form of light. The light is generated by microscopic devices called VCSELs. A VCSEL comprises three main components: a so-called quantum well and two mirrors. The quantum well turns some of the electricity that runs through the host chip into light, while the mirrors amplify it.
VCSELs are easier to manufacture than the lasers that optical networking devices typically use to generate laser. As a result, they often cost less. Another benefit of VCSELs is that they can be made from a compound called gallium arsenide. It’s more readily available than the materials that are most commonly used to produce miniature lasers.
Volantis will ship its chips with a data center inference appliance called the A-1. The system is about a third the size of a standard server rack. According to the company, it features 10 terabytes of memory with 250 terabits per second of memory bandwidth. Volantis estimates that the system will be capable of processing up to 10,000 tokens per second when running a model with 20 trillion parameters.
“This will enable real-time frontier inference, restart scaling laws & enable entire code bases in context windows,” Volantis co-founder and Chief Executive Officer Tapa Ghosh wrote in a blog post. “As a starting point, imagine a coding agent that completes a task in 30 seconds rather than 30 minutes.”
Volantis plans to start shipping its A-1 system next year.
