Reflection announced Beam on October 5, 2026, giving coding-agent builders and AI platform teams a new model to evaluate through a selective API beta. The company describes a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token.
Access is waitlisted, not generally available. Reflection plans Apache-2.0 weights, a technical report, model card and a fuller deployment package later in October; it has not named a release day. As of October 6, its official Hugging Face organization lists no public models.
What the beta supports
The documented model ID is Beam-501B-A23B. The following limits describe the hosted beta, not the eventual self-hosted release; Reflection says beta behavior and limits may change.
Area | Documented boundary |
|---|---|
API compatibility | Chat Completions and Models only; the Responses endpoint is not supported. |
Inputs and outputs | Text only, with streaming, tool calling and structured JSON output support. No native image, audio or file inputs. |
Token budget | 262,144 tokens for prompt and output combined; at most 131,072 generated tokens. |
Reasoning | Always on; accepts |
Budget calculation: reserving the full 131,072-token output allowance leaves at most 131,072 tokens for prompt-side content (262,144 − 131,072). Reasoning consumes that output allowance, so it is not a promise of 131,072 tokens of final-answer text. This is arithmetic from the documented limits, not a tested request. Reflection’s separate 1M-token midtraining claim is not the current API allowance.
An existing agent adapter also needs a compatibility check: stop strings are accepted but ignored and n must be 1. During tool rounds, preserve the returned assistant message, including tool_calls and reasoning_content.
One comparison trap is documented before any test begins: the API defaults to medium reasoning, while Mirror CLI defaults to xhigh. Set effort explicitly when comparing runs. Otherwise, using defaults can compare different configurations—a documentation-based inference, not a measured performance penalty.
What remains unverified
Reflection publishes coding and agent benchmark scores, but those are vendor-reported results, not independent reproduction. Its estimate of lower generation compute excludes prompt processing, context-dependent attention and serving overhead. It does not establish cheaper API bills or faster end-to-end runs.
The reviewed public documentation does not establish token prices or standard numeric quotas. Rate limits cover organization-wide requests and tokens, with concurrency controls and possible daily allowances. Confirm account pricing and capacity before budgeting a pilot; an unknown price is not a free service.
Evaluate now or wait?
Pilot after admission if the immediate question is text-based coding or tool-use fit. Pin the repository revision, agent harness version, reasoning effort and token budget. Define acceptance tests before running, then record failures, retries, usage and human acceptance.
Wait for the weights if the decision depends on local deployment, checkpoint inspection or reproducible self-hosted evaluation. Review the released license, model card and serving guidance before choosing hardware; the active-parameter count alone is not a deployment specification.
For a pilot involving code migration, the Mistral parity-harness guide explains how to define behavioral acceptance checks. It is a method for evaluation, not evidence of Beam’s performance.
Methodology: AI-assisted reporting and analysis based on Reflection’s announcement, developer documentation and public model listing, checked on October 6, 2026. No authenticated API calls, CLI runs or independent benchmark tests were performed.
