Article

llama.cpp /v1/systemone: Running Typed Decision Models Locally

Serve local routing and classification models with llama.cpp’s new endpoint, choose a checkpoint, and check compatibility before replacing a hosted API.

Editorial illustration for llama.cpp /v1/systemone: Running Typed Decision Models Locally: a geometric block represents a model release. Not documentary evidence.

On October 2, 2026, llama.cpp added /v1/systemone to llama-server, giving builders a local endpoint for supported decision models that return labels, scores and probabilities. The change shipped in b11361, a pre-release build. It is useful for routing requests, classifying documents and judging narrowly defined criteria without asking a chat model to write an answer.

The initial integration supports Laya, Julia-1, lev, OpenJev and Kev. b11364 added Bespoke Nimble support and was published on October 3 at 01:03 UTC, also as a pre-release. These are serving changes, not evidence that the models match a hosted service’s accuracy.

A local decision model can classify a support ticket into billing, technical or review before application code chooses the next step. That separation is the useful design pattern: the model interprets text; code retains business rules, permissions and fallback behavior.

Choose a checkpoint, not just an endpoint

The project announcement identifies the initial model sizes below. The linked checkpoint cards declare their weight licenses. llama.cpp’s MIT runtime license does not replace those terms.

Checkpoint

Size

Declared weight license

Local implementation boundary

Julia-1

144M

Apache-2.0

Small text checkpoint; use the author’s native option and context limits.

Laya

421M

Apache-2.0

The distributed checkpoint is English; do not assume family-wide multilingual routing.

Kev-4B

4B

Apache-2.0

Text; reference date preprocessing is not included.

lev

4B

Apache-2.0

Text decision model; evaluate its outputs against your labels.

OpenJev

27B

CC BY-NC 4.0

Image input needs its matching projector; weights are non-commercial.

Bespoke-Nimble-9B-v3

9B base

CC BY-NC 4.0 adapter

LoRA on Qwen3.5-9B; b11364 integration is text-only.

The b11364 converter explicitly omits Kev’s reference date preprocessing and marks Nimble image input unsupported. Julia-1’s author specifies 2–20 options and an 8,192-token combined limit. A runtime accepting a larger request is not proof of reliable classification at that size.

For a download-size reference, published Q8_0 files are approximately 168 MB for Julia-1, 449 MB for Laya and 4.48 GB for Kev-4B in decimal units. These are file sizes, not RAM or VRAM requirements; inference also needs working memory.

Clef is a separate version boundary. Its text-only integration merged on October 3, but the b11368 tag still predates that merge. Do not assume Clef support in b11364, or transfer hosted Clef’s vision capability to the new local integration.

Start one local text-routing service

For this example, obtain a platform-appropriate b11364 binary or build its source commit 46ca246de9bb1c35269722a6240d37d9dfd79cad. Check the build identity with llama-server --version. The integration PR lists pre-converted checkpoints; the following documentation-derived commands have not been executed for this article.

llama-server --host 127.0.0.1 --port 8080 -hf ggml-org/Kev-4B-GGUF:Q8_0

This keeps the example on loopback. The initial -hf download requires network access and is not an immutable model pin. Before deployment, record the binary hash, model repository revision, GGUF filename and hash, device/backend, context setting, question schema and thresholds. Load a retained GGUF with -m /path/to/model.gguf when reproducing that configuration.

Once the server is ready, save this synthetic request as request.json. Its state-and-questions structure asks for a route and a separate outage probability. It does not send the ticket or authorize an action.

{
  "state": {
    "ticket": "The export job stops with a timeout; invoices and payments are unaffected."
  },
  "questions": {
    "route": {
      "type": "choice",
      "instructions": "Choose the team supported by the ticket. Use review when the evidence is insufficient.",
      "criteria": {
        "billing": "Invoice, charge or payment issues",
        "technical": "Software faults and configuration problems",
        "review": "Unclear or outside these teams"
      }
    },
    "outage": {
      "type": "noul",
      "instructions": "Does the ticket explicitly report a service-wide outage?"
    }
  }
}
curl --fail-with-body --max-time 30 http://127.0.0.1:8080/v1/systemone \
  -H "Content-Type: application/json" \
  --data-binary @request.json

Validate the answer before using it

Read answers.route.choice and answers.route.probabilities. Read answers.outage.noul as a number between zero and one—not as the truthiness of the enclosing object. In the reference answer contract, score is a probability-weighted level index and can be fractional. The response will depend on the checkpoint and input.

The local endpoint documentation rules out streaming and reports zero output tokens. Input processing still costs work; zero generated tokens is not zero inference cost.

Use the numerical invariants in the upstream test source as integration checks, not as evidence of task accuracy. Before accepting an answer, verify:

  • The expected question IDs and answer types are present, and the selected route belongs to the allowed label set.

  • Probabilities are finite, within zero to one, and sum to one within an explicit rounding tolerance; the selected label agrees with a maximum.

  • A timeout, malformed answer, uncertain result or review label enters a defined fallback path rather than silently triggering a business action.

Error handling also changes. llama.cpp documents 400 for invalid requests and 501 for unsupported models or image configurations. TypeSafe’s hosted API instead documents 422 validation failures and 429/529 throttling or overload. Do not retry a capability mismatch as if it were temporary congestion.

A confidence threshold depends on the label set

The pinned choice implementation calculates confidence from the highest probability and the number of options, K, for two or more options:

confidence = (p_max - 1/K) / (1 - 1/K)

Illustrative calculation: hold the winning probability at 0.80. With two labels, confidence is 0.60; with four, it is about 0.733. Conversely, a confidence cutoff of 0.80 requires a winning probability of 0.90 for two labels but 0.85 for four.

This arithmetic isolates the formula; changing the real options can also change their probabilities. Version the label set with the acceptance policy. A confidence of 0.80 is not a measured 80% success rate, and a threshold fitted to a hosted model should not be carried over unchanged to a local checkpoint.

Evaluate the exact deployment before replacing a hosted route

A useful migration evaluation should answer whether the local configuration makes acceptable decisions on your workload—not merely whether it returns valid JSON. Proposed steps:

  • Freeze representative labeled requests, including rare categories, ambiguous text, unsupported requests, relevant languages and long inputs. Separate threshold tuning from the final evaluation set.

  • Compare the existing rules or hosted route with the exact local GGUF. Include near-threshold cases and changes to option order; record costly per-class mistakes and the fraction deferred for review.

  • Measure end-to-end p50/p95 latency at intended concurrency, memory use and cold starts separately. If testing several questions together, check whether neighboring questions change results; do not assume behavioral isolation across implementations.

  • Keep dependent decisions as explicit application steps. Treat a guardrail score as evidence for policy, not as permission to execute a tool or bypass a rule.

Quantization belongs in that evaluation. A smaller file is not evidence of unchanged probabilities or acceptance rates. The related GGUF-in-Transformers guide explains the broader local-model evaluation workflow; here, retain the decision schema and thresholds alongside the model artifact.

The project’s release article reports single-question GPU timings on an NVIDIA RTX PRO 6000. Those source-reported medians do not establish laptop performance, HTTP tail latency or equal accuracy among checkpoints.

Compare total cost, not just per-token fees

Published input rates checked on October 3 are $0.042 per million tokens for Jev 1.13, $0.09 for Clef-flash and $0.24 for Clef.

Illustrative arithmetic: one million decisions at 1,000 billable input tokens each equals one billion tokens, or $42, $90 and $240 respectively. This assumes equal token counts only for comparison, no retries and no free allocations or discounts. It does not assume equal model quality or actual tokenization.

Compare those bills with local hardware amortization, power, operations and fallback costs at the same acceptance target. Local inference may be worthwhile for data placement or an existing hardware budget without winning on price. For the managed alternative, see the Clef on Workers AI guide. Start with one bounded task; replace the hosted route only after the pinned local configuration meets its error, latency and cost requirements.

Methodology: AI-assisted reporting and implementation analysis based on GitHub releases, version-pinned source, model cards and provider documentation, checked October 3, 2026. Commands and evaluation steps are illustrative; no models or benchmarks were run for this guide.