Aleph Alpha released Kolibri-1 on October 3, 2026, giving teams building German- and English-language assistants an open-weight model they can operate themselves. The company lists it as generally available. Its practical appeal is control over inference for internal documents and workflows, particularly where an externally hosted endpoint is unsuitable.
Shortlist it if German and English dominate your workload, you have datacenter GPU capacity and an operations owner, and people can review the answers. Start with a document-grounded assistant. The release is not, by itself, evidence that Kolibri meets your accuracy, cost or regulatory requirements.
Choose the checkpoint before choosing the GPUs
Kolibri-1 is available in FP8 and BF16 checkpoints, with different hardware requirements. FP8 uses mostly 8-bit weights; BF16 uses 16-bit weights. The model card specifies 78.1 billion total parameters and about 3.46 billion active per token. The active count describes sparse computation, not the memory needed to hold the model.
The table separates Aleph Alpha’s FP8 guidance from its BF16 guidance. Each hardware entry is an alternative configuration, not a list of GPUs to combine. These are developer recommendations, not hardware tested for this article.
Deployment detail | FP8 | BF16 |
|---|---|---|
Checkpoint | Aleph-Alpha/Kolibri-1 | Aleph-Alpha/Kolibri-1-BF16 |
Vendor’s approximate weight footprint | 78 GB | 156 GB |
Vendor-listed minimum | 2× A100 80 GB; 2× H100 SXM5; 1× H200; 1× B200; or 1× B300 | 4× A100 80 GB; 4× H100 SXM5; 2× H200; 1× B200; or 1× B300 |
Vendor-recommended configurations | 2× H100 SXM5; 2× H200; 1× B200; or 1× B300 | 4× H100 SXM5; 2× H200; 2× B200; or 1× B300 |
Calculated parameter storage, not measured runtime memory: the FP8 repository metadata lists 77,398,016,000 one-byte parameters and 705,058,560 two-byte parameters. Multiplying and adding gives 78.81 GB, or 73.40 GiB. The BF16 metadata gives 78,103,074,560 × 2 bytes = 156.21 GB, or 145.48 GiB. GB here means decimal gigabytes; GiB uses powers of two.
These calculations exclude quantization scales, the key/value attention cache, activations and serving workspaces. FP8 nearly halves parameter storage, but that is not a measured halving of total GPU memory or cost. The published FP8 checkpoint’s parameters alone exceed 64 GB; a 24–64 GB device is not a straightforward fully resident deployment target.
Before ordering hardware, verify the exact accelerator, interconnect, supported kernels and parallelism with the intended runtime. Reserve capacity for simultaneous requests and generated output, not just loading the weights.
Treat 262,144 tokens as the starting ceiling
Aleph Alpha recommends at most 262,144 tokens for serving efficiency and complex tasks. The BF16 card distinguishes this trained context length from extension to 1,048,576 tokens, which the company says it has validated. A one-million-token headline is therefore not the default capacity or a guarantee of useful answers throughout that window.
For a document assistant, begin with smaller context tiers and increase only when a concrete question needs more evidence. Compare retrieving selected passages against sending the whole document collection. Keep prompt and generated output within the configured limit, and test whether answers remain supported when relevant text moves from the beginning to the middle or end.
Kolibri is text-in, text-out. Scanned forms, page layouts and charts need a separate conversion or visual-processing step. A long context window does not supply OCR.
Keep the serving environment versioned
The current aleph-alpha-inference 1.0.0 dependency file specifies vllm>=0.29.0,<0.30.0. Do not assume an existing vLLM 0.30 fleet can load this plugin unchanged. Use the supported environment, or wait for documented compatibility with your fleet version.
The official serving recipe uses --reasoning-parser kolibri1, --tool-call-parser kolibri1 and --enable-auto-tool-choice. Its FP8 recipe also uses --kv-cache-dtype fp8; the BF16 recipe changes the checkpoint and drops that cache flag. The short recipe is not a complete multi-GPU deployment configuration.
Reasoning is enabled by default. The model card exposes none, low, medium and high effort through chat_template_kwargs. Check that your client separates reasoning from answer content and correctly handles structured tool calls. Include the selected effort level in performance comparisons.
Pin the checkpoint, tokenizer, plugin, runtime, precision and deployment settings together so a rollback restores the same system. The RohitAI vLLM 0.30 guide explains model-specific serving profiles; it is not evidence that Kolibri supports 0.30.
Select by the task, not the strongest benchmark
Aleph Alpha’s launch comparison reports stronger German mathematics results for Kolibri than Qwen3.6-35B-A3B, but lower scores on tool calling and long-context work. Selected vendor-reported scores follow; higher is better. They are not independently reproduced results or success rates for your application.
Benchmark | Kolibri | Qwen3.6-35B-A3B |
|---|---|---|
AIME 2026, German mathematics | 90.0 | 84.4 |
BFCL v4 overall, function calling | 61.4 | 67.2 |
LongBench Pro, long-context tasks | 64.5 | 70.8 |
For an internal German knowledge assistant, the decisive questions are whether answers cite the right evidence, whether unsupported questions trigger abstention, and whether tool arguments remain correct across several turns. A mathematics score does not answer those questions.
Separate open weights from deployment assurances
The model card’s license statement applies Apache 2.0 to the published weights and configuration files, not to unreleased training artifacts. The inference plugin has its own Apache-2.0 license. This is an open-weight release, not publication of the complete training stack.
The repository’s Apache license includes redistribution conditions and warranty disclaimers. A license grant is not a compliance certification or a support contract. For a regulated workflow, map where inference, retrieval, logs, backups and external tool calls run; local weights alone do not establish that all data stays inside your environment.
The reviewed official product material did not establish a public Kolibri API token tariff, hosting SLA or enterprise-support price. Obtain a scoped quote if those services matter. Self-hosting still requires GPU capacity, engineering and ongoing operations.
A bounded evaluation before deployment
The following is a proposed evaluation plan, not a report of tests performed. Define acceptance criteria with the people who own the workflow before comparing Kolibri with the currently approved baseline.
Check bilingual evidence handling. Use approved, held-out German and English documents with expert-reviewed answers. Include administrative terminology, conflicting document versions and questions whose answers are absent. Score citation-supported correctness, false assertions, appropriate abstention and unnecessary refusals separately.
Vary context and precision. Start with 32k, 128k and 256k context tiers, keeping space for output. Compare BF16 and FP8 on identical tasks and quality thresholds. Extend beyond 262,144 tokens only if a specific task benefits enough to justify the extra capacity and latency.
Exercise tools without granting broad authority. Check tool names, argument schemas, multi-turn recovery and duplicate side effects. Include malicious instructions inside retrieved documents. Enforce permissions and idempotency in application code; require human approval for consequential actions. A parsed call is not permission to execute it.
Measure the service people will use. Record time to first token, median and p95 completion latency, peak GPU memory, timeouts and concurrency for the chosen reasoning settings. Count an accepted task only when it meets the agreed quality standard. Compare total operating cost per accepted task, including review and retries, at the required response time.
Proceed with a limited document-assistant pilot if Kolibri clears those checks and private inference solves a real requirement. If it fails evidence handling or cannot meet latency within the available GPU budget, retain the baseline. German-English specialization and downloadable weights justify evaluation; the workflow results should decide deployment.
Methodology: AI-assisted reporting and analysis based on Aleph Alpha’s announcement, pinned model cards and inference-package source, plus Hugging Face metadata, checked on October 3, 2026. No model execution or performance testing was conducted for this article. Memory figures labeled as calculations are arithmetic estimates, and benchmark scores are Aleph Alpha’s claims.
