Reflection AI Inc. today introduced Beam, an open-source large language model with 501 billion parameters.
The launch comes a few months after the startup raised funding at a $25 billion valuation. Around the same time, Reflection AI reportedly inked a $6.3 billion deal with SpaceX Corp. to rent Nvidia GB300 NVL72 appliances. It used those systems, which each contain 72 graphics cards, to train Beam.
Reflection AI benchmarked its model against GLM-5.2, an open-source LLM that includes about 250 billion more parameters. The company determined that Beam can perform some tasks better while using between one third and one fourth the hardware. Furthermore, it says that the LLM approaches the performance of Qwen 3.8-Max, a model with over 2 trillion parameters.
The launch is particularly notable because many of the open-source ecosystem’s most advanced LLMs were created by Chinese companies. That includes Qwen 3.8-Max and GLM-5.2. Beam is the first open-source model from a U.S. startup to have demonstrated comparable or better performance. However, free LLMs continue to trail frontier models such as Anthropic PBC’s Claude Fable 5.1.
Reflection AI kicked off the development of Beam by training a relatively small prototype model. It then created a series of successively larger, more capable algorithms. The workflow eventually produced Beam Base, the foundation on which Beam is based.
Reflection AI developed Beam Base using a cluster of 6,144 graphics cards. It trained the model on 23.8 trillion tokens sourced from the public web and commercial sources. According to Reflection AI, that dataset included a significant amount of software code. The company created custom filters for each programming language to remove low-quality files.
Reflection AI developed Beam Base in under four weeks. It subsequently performed midtraining, an optimization process that extended the model’s context window and enhanced its reasoning capabilities. That phase of the project set the stage for the third, most hardware-intensive phase of the training workflow.
Reflection AI spun up 10,000 GB300 graphics cards and used them to launch 1.3 billion reinforcement learning sandboxes. Those are virtual environments in which an AI model learns new skills. Beam’s sandboxes were optimized for tasks such as generating code, searching the web and running AI agents.
Reflection AI says that the project’s reinforcement learning phase took only four weeks. The company avoided unnecessary delays by developing software that prevented malfunctions from interrupting the workflow. According to Reflection AI, its cluster achieved a median recovery time of eight minutes across the 71 errors that cropped up during training.
Beam is initially available through an early access program. Reflection AI plans to release its weights, documentation and fine-tuning tools later this month.
