Reflection’s Beam model hits 501 billion parameters

Reflection has introduced Beam, a 501 billion parameter sparse Mixture-of-Experts model trained to excel at coding, reasoning, and agentic tasks. The model’s development involved a substantial investment in computational resources; a four-week reinforcement learning run used 10.5K NVIDIA GB300 GPUs to generate over 100 million rollouts. Beam was pretrained on 23.8 trillion tokens, achieving performance competitive with GLM 5.2 and approaching Qwen 3.8-Max on key tasks while offering efficiency advantages over Kimi K3. These results, the company states, translate into “more intelligence per token,” delivering strong capabilities at a lower cost.

Beam: Reflection’s 501B Parameter Sparse Mixture-of-Experts Model

Reflection detailed a novel approach to evaluating model performance, moving beyond simple parameter counts to assess generalization ability through an experiment. The team tasked Beam with generating a grid representing longitudes and latitudes, a puzzle appearing only days before the evaluation and therefore absent from the training data; Beam achieved 95.5% accuracy, positioning it between the Opus 5 and Fable 5 models in this specific test of novel task completion, Reflection says.

This assessment method, focusing on out-of-sample generalization, provides a more detailed understanding of Beam’s capabilities than raw parameter size alone. The development of Beam incorporated a multi-tiered safety and alignment strategy, prioritizing adherence to defined principles during reinforcement learning. Researchers organized these principles into three levels, beginning with foundational rules the model should never violate, such as upholding safety policies and maintaining its identity as an AI agent.

This structured approach to alignment aimed to predictably shape Beam’s behavior through incentivized environments and generative reward models, allowing for more controlled development of the model’s responses. They found they could predict reinforcement learning gains with a correlation coefficient of 0.79, exceeding the predictive power of a Best-of-N ceiling alone at 0.46.

The computational scale required to train Beam was substantial, yet the company focused on achieving efficiency in inference, according to Reflection. Using estimates based on active parameter count and generated tokens, Reflection calculated that Beam’s forward-pass compute demands are lower than those of larger models like Qwen 3.8-Max, which require significantly more computational resources per token.

Pretraining Beam involved a curated dataset of 23.8 trillion tokens, sourced from both publicly available web data and proprietary licensed datasets. This extensive dataset was designed to match or surpass the quality and scale of those used to train comparable open-weight models, providing a strong foundation for Beam’s coding, reasoning and agentic capabilities. The selection of high-quality tokens was a key focus, ensuring the model learned from reliable and informative sources.

Reflection plans to release Beam’s weights under an Apache 2.0 license this month, alongside comprehensive documentation and tools for running, evaluating, and fine-tuning the model. This commitment to open access aims to encourage wider adoption and collaborative development within the AI community, allowing developers to integrate Beam into existing open-source workflows. The company is also partnering with distribution partners to facilitate broader accessibility and integration.

Beam represents the first in a series of models Reflection intends to release, signaling a long-term commitment to advancing open intelligence. The team is already working on subsequent iterations, with the goal of continually pushing the boundaries of AI capabilities and making the latest advancements available to a wider audience. “We are already training what comes next,” the company stated, “with the goal of bringing the open frontier closer to the frontier of intelligence with every release.”

Pretraining Beam on 23.8 Trillion Tokens for Coding and Reasoning

Beam’s pretraining used 23.8 trillion tokens, a curated collection designed to bolster coding and agentic performance beyond simply scaling model size. This dataset prioritized source code, technical explanations and mathematical/scientific knowledge, undergoing a tiered quality assessment to weight training toward the most reliable material. Reflection developed internal classifiers to evaluate web content, code and STEM data, ensuring a foundation suited for complex downstream tasks. The pipeline’s focus extended to nearly all publicly accessible and unrestrictively-licensed code documentation available on the web, establishing a robust base for agentic coding applications.

The emphasis on data quality directly informed the model’s architecture and training regimen, enabling longer interactions and more demanding environments during reinforcement learning. This approach facilitated throughput, failure recovery and integrity checks throughout the learning process, ultimately contributing to Beam’s capabilities in reasoning and tool use.

While the model itself is text-only, it can process information from other modalities when presented as text, demonstrating versatility across a range of applications. For example, Beam successfully generated a fine-tuning notebook for the Gemma-4 model on a Text2SQL task, highlighting its adaptability to diverse challenges. Beam’s performance on advanced reasoning benchmarks achieves scores comparable to GLM-5.2 while requiring three to four times less inference compute.

This efficiency advantage positions Beam as a competitive alternative to larger open models like Kimi K3, which currently lead in raw capability. Evaluation across coding, agentic, reasoning, and STEM benchmarks demonstrates Beam’s breadth, with results on the Terminal Bench v2.1 reaching 80.1, and SWE Bench Pro v2-Hard scoring 77.2. The team also reports a score of 79.3 on the AA-LCR benchmark, and 65.5 on LongBench v2, showing its ability to handle complex tasks and extended contexts.

The development of Beam involved an iterative process of pretraining progressively larger models to refine and validate the scaling strategy. This approach allowed Reflection to establish a foundation with rich coding knowledge and inherent agentic capabilities, amplifying its potential through stable Mixture-of-Experts optimization, the firm reports. Demonstrations include building a live NYC subway dashboard using public data and creating interactive applications, showing Beam’s ability to combine reasoning, coding and tool use in practical scenarios.

The curated dataset and efficient architecture contribute to Beam’s ability to tackle complex tasks while maintaining performance and stability. The company’s focus on quality-centric data curation and robust reinforcement learning techniques positions Beam as a significant advancement in open-weight models, offering a balance between capability and efficiency.

High-Compute Reinforcement Learning with 10.5K NVIDIA GB300 GPUs

The training of Reflection’s Beam model used 10.5K NVIDIA GB300 GPUs. This substantial investment in compute infrastructure enabled the exploration of more extensive problem-solving strategies, extending rollout lengths to support multi-step reasoning and adaptation to environmental feedback. The team designed this approach to translate increased computational power directly into improved model capabilities, a focus reflected in the scale of the hardware deployment.

Scaling reinforcement learning required not only hardware but also a robust environment for training; Reflection constructed a pool of nearly one million environments, primarily through synthetic data pipelines, supplemented by proprietary and open-source data. These environments were designed to present difficult, high-quality tasks, allowing the model to train on longer interactions and more demanding scenarios while maintaining throughput and ensuring the integrity of the learning process. The creation of this extensive training ground emphasises the importance of task diversity in achieving robust and adaptable artificial intelligence.

Beam’s foundation was established through a series of progressively larger pretrained models, a process designed to verify a specific scaling recipe and ensure stable Mixture-of-Experts optimization dynamics. This recipe incorporated techniques like SandwichNorm, elementwise attention gating and FP32 residual accumulation to control activation growth and minimize rounding errors, thereby maintaining healthy signal propagation throughout the model’s layers. Residual stream RMS remained bounded throughout pretraining across all 52 layers of Beam, exhibiting smooth, depth-dependent trajectories and avoiding sustained activation growth, according to the company.

Careful attention was paid to data curation, including repeated exposure to code and technical content, employing fuzzy deduplication and packing algorithms to maximize generalization. “We repeated code and technical content multiple times to increase the model’s exposure over the course of its training horizon,” the team stated, emphasizing the importance of strategic data repetition.

A dedicated midtraining stage was implemented to specifically prepare Beam for high-compute reinforcement learning, developing knowledge, reasoning and tool-use capabilities. Multi-stage data curation pipelines were constructed to capture the complexity of real-world tasks, extending coverage of capabilities difficult to learn from raw data alone. Starting with carefully selected real-world examples, these pipelines transformed, combined, and extended material into training data designed to teach specific skills.

This approach suggests a deliberate effort to move beyond simple data scaling and towards more targeted, curriculum-based learning. The team’s focus on building a strong prior for reinforcement learning is evident in their infrastructure choices and data curation strategies. By investing in both the hardware and the software necessary to support large-scale RL, Reflection aims to push the boundaries of what is possible with open-weight models and contribute to a more accessible and collaborative AI ecosystem.

The combination of high-compute infrastructure, curated datasets and advanced training techniques positions Beam as a significant step forward in the development of capable and efficient AI systems, the company states. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training.

Beam’s Coding and Agentic Performance Versus Open-Weight Models

Beam achieves performance competitive with GLM 5.2 and approaches Qwen 3.8-Max on coding and agentic tasks, according to Reflection’s data, while offering an advantage in inference efficiency. The 501 billion parameter model demonstrates this capability across a range of benchmarks, including SWE Bench Pro v2-Hard where it scored 77.2, and Terminal Bench v2.1 where it reached 80.1. These results position Beam as a viable alternative to larger models requiring more computational resources for similar performance.

The model’s efficiency gains are particularly noticeable when compared to models like Kimi K3, which, while exhibiting higher raw capability, demands more inference compute per token. The company highlights this as a key differentiator, suggesting a focus on practical deployment alongside performance, the company’s account states. Reinforcement learning played an important role in optimizing Beam’s performance, initially improving task completion with reduced reasoning steps.

As the model developed stronger agentic capabilities, completion lengths increased, but these additional tokens supported further gains in performance, indicating a refined balance between capability and resource usage. “Early in RL, performance improved even as completion lengths fell: the model learned to solve tasks more effectively with less reasoning,” the team reports. The data reveals a two-phase progression during the reinforcement learning run, as demonstrated by DeepSWE scores.

Initially, the Pareto frontier moved towards increased reasoning effort, and later, as agentic capabilities strengthened, the frontier shifted again. This suggests a deliberate strategy to first optimize for efficiency and then expand the model’s capacity for complex problem-solving. Beam’s performance on AA-LCR reached 79.3, while it scored 65.5 on LongBench v2, demonstrating its ability to handle both short-form and extended reasoning tasks.

Reflection used 23.8 trillion tokens for Beam’s pretraining, a curated collection intended to match or surpass similar-sized open base models in quality and diversity, the company claims. This extensive dataset, combined with a focus on coding and agentic performance, appears to have yielded a model capable of competing with established players in the open-weight space. The company’s approach emphasizes not just scale, but also the careful selection and curation of training data to maximize the model’s potential.

The model’s architecture, a sparse Mixture-of-Experts design, contributes to its efficiency and scalability, Reflection says. This allows Beam to activate only a subset of its parameters for each task, reducing computational demands without sacrificing performance. The team’s focus on stable MoE optimization dynamics suggests a commitment to building a robust and reliable model capable of handling complex workloads. Beam’s capabilities, as demonstrated by its performance on benchmarks like SWEAtlas Codebase QnA (scoring 34.6), position it as a powerful tool for enterprise coding and agentic applications, offering a balance of performance, efficiency and cost-effectiveness.

Asynchronous Policy Gradients Stabilize Large-Scale RL Training

Training Beam at scale necessitated innovations in reinforcement learning stability, moving beyond 30,000 rollouts used for Inkling and the 753,000 rollouts of MiMo. The team employed asynchronous policy gradients, but recognized that at this magnitude, policy staleness, the delay between data generation and model updates, posed a significant challenge, according to Reflection. Completed rollouts often originated from multiple model checkpoints, with older tokens becoming increasingly misaligned with the current policy.

Numerical discrepancies between training and inference engines further exacerbated this issue. To counter these effects, developers created new algorithms designed to maintain stable learning even with stale data, while simultaneously reducing training-inference mismatch throughout the entire pipeline. These advances enabled fully asynchronous reinforcement learning, sustaining stability even when learning from interactions generated more than 24 hours earlier. Analysis revealed that learning remained consistent as staleness increased, with the oldest samples in training batches demonstrating stable numerics.

Even with a 107-version lag between the current policy and the training data, the numerical stability remained intact. Controlling reasoning length proved important to Beam’s development, prompting the implementation of a controllable length penalty during training. This system rewards successful solutions while discouraging superfluous tokens, initially contracting the policy to prioritize token efficiency. Subsequently, the model expanded its reasoning efforts to achieve high performance, allowing users to adjust the balance between response length and accuracy, the company says.

This flexibility enables users to tailor reasoning effort to specific tasks and available compute resources. Beam’s reinforcement learning training was specifically designed to support reasoning and agentic capabilities that generalize beyond the initial training tasks. Throughout development, the team found that data quality was paramount, with compromises leading to capability plateaus. Systematic improvements to task quality were essential for sustaining gains throughout the training run, which showed no signs of saturation as it concluded.

Supporting this high-compute agentic reinforcement learning required a platform capable of generating rollouts, executing tools, evaluating outcomes and updating the model at scale. The team built an asynchronous platform allowing these processes to run independently, coordinating the flow of experience and model updates. During Beam’s training, they sustained an average of 110,000 concurrent rollouts, facilitated by seven key capabilities. These included fully asynchronous execution, where agents generate rollouts while the trainer learns and publishes new model versions, tagging each token with its originating version to account for policy staleness.

Flexible compute allocation also played a major role, with the team adjusting the balance between inference and training as the workload evolved, operating at inference-to-training GPU ratios ranging from 3.9:1 to 5.4:1. Safety and alignment were addressed through a separate training pipeline, using a dedicated SFT and reinforcement learning approach on data specifically designed to instill desired principles.

The capabilities of this safety teacher were then merged with those of a large-scale RL teacher via multi-teacher on-policy distillation. This allowed for iterative refinement of rubrics and reward design, suppressing undesirable behaviors like hallucinations and excessive formatting, and improving overall interaction quality.

Their safety training used deliberative alignment techniques to directly incorporate the safety policy into the model’s reasoning process, building the dataset adversarially and iteratively. In each round, they trained a model, generated prompts that elicited harmful or overly cautious behavior, and incorporated those successful attacks into the SFT mixture for the next iteration.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Rusty Flint

Rusty Flint

Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things Quantum Physics especially. Rusty likes the more esoteric side of Quantum Computing and the Quantum world. Everything from Quantum Entanglement to Quantum Physics. Rusty thinks that we are in the 1950s quantum equivalent of the classical computing world. While other quantum journalists focus on IBM's latest chip or which startup just raised $50 million, Rusty's over here writing 3,000-word deep dives on whether quantum entanglement might explain why you sometimes think about someone right before they text you. (Spoiler: it doesn't, but the exploration is fascinating)

Latest Posts by Rusty Flint: