Reflection announced Beam on October 5 — its first open-weight model, and one of the few "Western" open releases that targets agentic coding head-on. The architecture is the headline: a sparse Mixture-of-Experts with 501 billion total parameters and 23 billion active. That is the detail to start with, because it decides whether the model fits your hardware at all.
Only 23B parameters participate in the forward pass per token, so the compute cost per token lands in the range of a mid-size dense model, while the knowledge lives in the full 501B. If you have ever priced a 70B dense model at Q4 against a MoE of similar total size, you already know the tradeoff. You still need memory for the whole model, but you generate at MoE-active speed.
The numbers that matter for build-vs-buy
Reflection reports the following (comparison models from their own table):
| Benchmark | Beam (501B, 23B active) | GLM 5.3 | Kimi K3 | Qwen 3.8-Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|
| SWE Bench Verified | 80.9 | NR | NR | NR | NR |
| SWE Bench Pro v2-Hard | 77.2 | 84.3 | 88.2 | NR | NR |
| Terminal Bench 2.1 | 80.1 | 88.2 | 88.3 | 86.6 | 90.6 |
| DeepSWE v1.1 | 44.4 | 61.0 | 68.0 | 51.0 | 74.2 |
| MCP Atlas | 78.7 | 84.2 | 82.3 | 84.5 | NR |
| GPQA Diamond | 90.5 | 91.7 | 93.5 | 92.6 | 90.9 |
| AIME 2026 | 97.8 | NR | NR | NR | NR |
NR means the model was not reported on that benchmark — several of those cells are missing for real, not zero. Read the table with that caveat: it is not a leaderboard, it is a subset of benchmarks Reflection chose, with comparison numbers they sourced from Artificial Analysis and DataCurve.
Where Beam does not win outright, the pitch is efficiency. Reflection claims performance "comparable to GLM-5.2 while using 3–4× less inference compute" on advanced reasoning benchmarks, and larger gains against 2T+ models like Qwen 3.8-Max. Figure 2 in the post is their own FLOP estimate — 2 × active parameters × mean generated tokens per attempt, counting each multiply-add as two operations — not a measured serving cost. Prefill, attention and serving overhead are excluded. Treat it as an order-of-magnitude argument, not a bill.
What you get to control at runtime
Beam exposes a reasoning effort parameter. Lower settings favor shorter responses, higher settings allow longer chains on hard tasks. This was trained in deliberately: Reflection applied a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens, and the post's Figure 5 shows the Pareto frontier contracting (more token-efficient) before the high-effort end expands outward. That second phase is the interesting part — after the model got efficient, extra reasoning tokens bought real performance gains rather than just filling the context window.
For agentic work, this is what I would actually tune. A code-review loop and a full SWE-bench-style task should not use the same effort setting. A knob that trades tokens for accuracy is more useful than a single "thinking budget" you have to guess at.
The RL run is the real story
The training numbers are unusually specific:
- Pretraining: 23.8 trillion tokens, run end-to-end in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs, finishing at 92.3% goodput with nine semi-automatic rewinds.
- RL: 10.5K GB300 GPUs for four weeks, generating over 100 million rollouts with a maximum context length of 256K tokens, and roughly 1.3 billion sandboxes. Reflection says this is one of the largest open-lab RL runs done to date.
- RL infrastructure sustained an average of 110K concurrent rollouts, scaling to 170K concurrent sandboxes across 20 clusters, two clouds and four regions. 90% of new sandboxes were ready in under 10 seconds.
- Inference-to-training GPU ratios were adjusted between 3.9:1 and 5.4:1 mid-run, and the trainer was resized across five GPU mesh configurations without losing state.
- New weights reached the inference fleet in a median of about 12 seconds; hierarchical distribution (RoCE across racks, NVLink locally) cut cross-rack traffic by 75% and sped fleet-wide adoption 2.2×.
- They trained with asynchronous policy gradients and report stable numerics even at one-day staleness — 107 weight versions behind the current policy.
For anyone running multi-agent systems, the midtraining detail is worth stealing: Beam's effective context was extended to 1M tokens specifically so RL could work over long horizons. Longer contexts are not just a retrieval feature — they are a training prerequisite for agentic tasks.
What to watch before you commit
The weights are not out yet. Reflection says it will release the weights under Apache 2.0, plus a technical report, model card and the tooling to run, evaluate and fine-tune, later this month. Today there is only a waitlist signup. Beam is also text-only — the post is explicit that it reasons about visuals as text, though it can call OCR APIs when given web access.
The catch for a self-hosting shop is the split you always get with MoE: memory has to hold the full 501B (even heavily quantized, that is a serious footprint), while throughput tracks the 23B active set. That is good for cost per token on a multi-GPU box, and bad if you were hoping for a single 24GB card.
I would watch three things when the weights ship: what the actual vLLM/SGLang integration looks like, whether reasoning effort is exposed as a stable API parameter or a harness-specific flag, and whether the safety evaluations land in the technical report alongside the capability numbers.
Sources: