Article

Security-One 27B: When to Use It for Agent Security Triage

Evaluate Security-One 27B for agent security triage: reported detection tradeoffs, API pricing, self-hosting requirements and a review-only pilot plan.

Editorial illustration for Security-One 27B: When to Use It for Agent Security Triage: a document and lock represent artifact protection. Not documentary evidence.

Superagent announced Security-One 27B on October 4, 2026, giving agent builders and security teams a model that scores security events rather than writing chat responses. It is offered through a hosted API and Apache-2.0 weights. The practical question is whether its first-pass classifications can help teams reserve deeper investigation for the events that need it.

The announcement is new; the checkpoint is older. Superagent’s model documentation records a September 25 release date. October 4 should not be read as the date of a new checkpoint update.

The strongest starting point is a review-only prompt-injection pilot. Superagent identifies that as its best-validated security capability; tool-call, code and infrastructure triage need their own evaluations. Use Security-One to route suspicious inputs for review, with permission checks outside the model. A favorable score must not override a denied tool, destination or privilege. The model card makes that separation explicit.

What the classifier returns

You supply the event or document and the questions to ask about it. The hosted API accepts authenticated POST /v1/systemone requests using model security-one. Its three answer types serve different decisions:

  • noul: the probability that a binary criterion applies, such as whether text attempts prompt injection.

  • choice: the highest-probability category plus the distribution over all supplied options.

  • score: the probability-weighted average of zero-based severity levels—not necessarily the most likely level.

None is an explanation of a vulnerability. For code review, use classification to decide what deserves investigation; a suspicious score does not establish a line-level finding.

One integration trap is treating confidence as the probability of an attack. For choice and score answers, the published scoring implementation defines confidence as one minus normalized entropy: how concentrated the distribution is. Calculated example: probabilities of 0.70 and 0.30 give confidence of about 0.119. A rule testing confidence ≥ 0.70 is therefore different from one testing P(unsafe) ≥ 0.70. Neither establishes correctness on your traffic.

Read the attack results alongside the false alarms

Superagent’s version-pinned evaluation report uses an unsafe-probability threshold of 0.70. These are developer-reported results, not independent measurements. The counts below are reconstructed from its dataset sizes and reported fractions.

Evaluation

Attacks detected

Benign inputs flagged

BIPIA binary adaptation

599 / 600 (99.83%)

1 / 200 (0.50%)

Deepset test set

47 / 60 (78.33%)

0 / 56 (0%)

NotInject

Not measured: benign-only

42 / 339 (12.39%)

BIPIA here is a binary detector adaptation, not the original generative benchmark. The NotInject researchers built a stress test of benign text containing attack-associated trigger words. Its result is a reason to include legitimate security discussions and quoted attacks in a pilot—not a forecast that 12.39% of ordinary traffic will be blocked. Likewise, zero false alarms in 56 examples does not prove a zero real-world rate.

The same vendor report shows a tradeoff against hosted Jev 1.13.0: Security-One catches more BIPIA and Deepset attacks at the chosen threshold, but flags more NotInject inputs. That is not a comparison at matched false-positive rates or evidence that one model is universally better. See the evaluation protocol and comparator results.

Illustrative queue calculation, not a traffic forecast: suppose 100,000 events contain 0.1% attacks, and the BIPIA rates transfer unchanged. Expected true flags are 100 × 599/600 ≈ 100; false flags are 99,900 × 1/200 ≈ 500. Only about 16.7% of flagged events would be attacks. These assumptions are deliberately hypothetical: the benchmark does not establish your attack prevalence or error rates.

That is why a pilot needs both missed-attack measurements and a review-capacity budget. A high detection percentage alone cannot tell you how many investigations the model will create.

Hosted API or self-hosted weights?

As checked on October 4, the public model document advertises security-one as ready. This confirms the published offering, not a successful authenticated inference call. The Hugging Face repository lists BF16 weights under Apache 2.0, but downloading them requires accepting its access conditions and sharing contact information.

Route

Best reason to consider it

What still needs validation

Hosted API

Pilot without operating a GPU service, if the selected data can leave your infrastructure.

Account entitlement, data terms, workload latency and actual billing.

Self-hosted weights

Keep inference inside infrastructure you control.

GPU capacity, serving operations and probability calibration.

Review-only or defer automation

You lack representative labels or a dependable escalation path.

Whether measured errors and review volume justify automated action.

These are deployment choices, not a measured cost ranking. The documented self-hosting configuration uses one NVIDIA B200, SGLang 0.5.19 and BF16 weights. Its validated classification context is 65,536 tokens; the model’s larger native window is not an equivalent validation claim.

For sizing intuition only, 27 billion parameters × two bytes per BF16 parameter is about 54 GB of weights before cache and runtime overhead. This calculation is not a minimum-VRAM specification or proof that a particular GPU fits the workload.

Self-hosting also requires the prescribed prompt construction, option-token readout and calibration temperature. Ordinary chat prompting is not equivalent. Re-evaluate after changing precision or runtime. RohitAI’s llama.cpp typed-decision serving guide provides broader local-serving context, not confirmation that Security-One works interchangeably with that endpoint.

Pricing: keep dollars and credits separate

The public model metadata lists input at USD 0.00000005 per token: USD 0.05 per million. It lists no output-token charge. The commercial pricing page instead quotes 0.05 credits per million Security-One input tokens.

The monthly platform plans show 100 credits free, 500 credits for USD 499, 2,000 for USD 1,499 and a 6,000-credit tier starting at USD 3,999. These are shared platform subscriptions, not dedicated Security-One hosting prices. Do not assume that one credit equals one dollar or that the metadata guarantees standalone pay-as-you-go access.

Illustrative usage calculation: one million decisions averaging 1,000 total billable input tokens each equals one billion tokens. That yields USD 50 under the provider metadata, or 50 credits under the credit schedule. Those are separate descriptions, not interchangeable bills. The example excludes retries, discounts, subscriptions, escalation and human review, and assumes any question overhead is already inside the token budget.

Confirm account-specific terms before estimating spend. Then measure the whole cascade: classifier usage plus larger-model calls and review work. For a general routing comparison, see Clef on Workers AI and its cascade economics.

What a useful pilot should establish

Start in shadow mode: record the proposed routing decision without letting it authorize an action. Superagent’s prompt-injection example illustrates allowing below 0.30, escalating from 0.30 to below 0.70, and blocking at 0.70 or above. Treat those bands as hypotheses to evaluate, not production defaults.

  • Separate tuning from evaluation. Build labeled samples of attacks and benign events, including security documentation, long inputs and relevant languages. Hold back data for a final check after threshold tuning. Published evaluation is primarily English, according to the model limitations.

  • Measure the review burden. Track precision, recall, consequential misses, escalation rate, latency under expected concurrency and total cost per decision. Evaluate tool-call and code triage separately from prompt injection.

  • Define missing-verdict behavior. Timeouts, invalid answers and exhausted retries must send consequential actions to a hold or review path, not silently permit them. Keep both the classifier and the reviewing model behind application-owned permission checks.

  • Make results reproducible. Record model revision, question and label definitions, serving configuration and policy thresholds together. Retest after changing any of them.

Plan for the documented API limits: 65,536 tokens per question including shared state, 131,072 total input tokens per request, and a 120-second deadline. Over-limit inputs are rejected rather than silently shortened. These are ceilings, not latency promises; bounded retries and explicit handling of unclassified content belong in the pilot.

Adopt automated routing only when held-out results show acceptable misses, false alarms and review volume for that specific workflow. Until then, Security-One is a candidate for prioritizing investigation, not a replacement for authorization or security review.

Methodology: AI-assisted reporting and analysis of Superagent’s announcement, documentation, model card and version-pinned evaluation data, checked October 4, 2026. No model inference, hardware benchmark or customer deployment was performed. Calculations are hypothetical; all reported model results are Superagent’s.