Article

Reflection Beam Enters Waitlisted Beta; Open Weights Planned for October

Reflection’s Beam beta supports text-based coding and agents, but Apache-2.0 weights are pending. What builders can evaluate now—and what must wait.

Editorial illustration for Reflection Beam Enters Waitlisted Beta; Open Weights Planned for October: a controlled task flows from input to output. Not documentary evidence.

Reflection announced Beam on October 5, 2026, giving coding-agent builders and AI platform teams a new model to evaluate through a selective API beta. The company describes a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token.

Access is waitlisted, not generally available. Reflection plans Apache-2.0 weights, a technical report, model card and a fuller deployment package later in October; it has not named a release day. As of October 6, its official Hugging Face organization lists no public models.

What the beta supports

The documented model ID is Beam-501B-A23B. The following limits describe the hosted beta, not the eventual self-hosted release; Reflection says beta behavior and limits may change.

Area

Documented boundary

API compatibility

Chat Completions and Models only; the Responses endpoint is not supported.

Inputs and outputs

Text only, with streaming, tool calling and structured JSON output support. No native image, audio or file inputs.

Token budget

262,144 tokens for prompt and output combined; at most 131,072 generated tokens.

Reasoning

Always on; accepts low, medium, high, xhigh and max. The API default is medium.

Budget calculation: reserving the full 131,072-token output allowance leaves at most 131,072 tokens for prompt-side content (262,144 − 131,072). Reasoning consumes that output allowance, so it is not a promise of 131,072 tokens of final-answer text. This is arithmetic from the documented limits, not a tested request. Reflection’s separate 1M-token midtraining claim is not the current API allowance.

An existing agent adapter also needs a compatibility check: stop strings are accepted but ignored and n must be 1. During tool rounds, preserve the returned assistant message, including tool_calls and reasoning_content.

One comparison trap is documented before any test begins: the API defaults to medium reasoning, while Mirror CLI defaults to xhigh. Set effort explicitly when comparing runs. Otherwise, using defaults can compare different configurations—a documentation-based inference, not a measured performance penalty.

What remains unverified

Reflection publishes coding and agent benchmark scores, but those are vendor-reported results, not independent reproduction. Its estimate of lower generation compute excludes prompt processing, context-dependent attention and serving overhead. It does not establish cheaper API bills or faster end-to-end runs.

The reviewed public documentation does not establish token prices or standard numeric quotas. Rate limits cover organization-wide requests and tokens, with concurrency controls and possible daily allowances. Confirm account pricing and capacity before budgeting a pilot; an unknown price is not a free service.

Evaluate now or wait?

  • Pilot after admission if the immediate question is text-based coding or tool-use fit. Pin the repository revision, agent harness version, reasoning effort and token budget. Define acceptance tests before running, then record failures, retries, usage and human acceptance.

  • Wait for the weights if the decision depends on local deployment, checkpoint inspection or reproducible self-hosted evaluation. Review the released license, model card and serving guidance before choosing hardware; the active-parameter count alone is not a deployment specification.

For a pilot involving code migration, the Mistral parity-harness guide explains how to define behavioral acceptance checks. It is a method for evaluation, not evidence of Beam’s performance.

Methodology: AI-assisted reporting and analysis based on Reflection’s announcement, developer documentation and public model listing, checked on October 6, 2026. No authenticated API calls, CLI runs or independent benchmark tests were performed.