LiteLLM’s October 11, 2026 stable release adds configurable message screening for teams running customer chatbots or internal assistants through its gateway—the software connecting their applications to AI models. Administrators can write yes/no screening questions and choose whether a flagged result is logged or blocks the request.
A separate decision model assesses the text against those questions; it does not write the chatbot’s answer. The new administration workflow lets teams select a model, set thresholds and try sample inputs. The practical choice is how much missed abuse or wrongly blocked legitimate traffic the business can tolerate.
Our recommendation: begin a new or ambiguous policy with logging and representative evaluation. Move to blocking only when the rule is clear and its mistakes are acceptable for that particular service. Decide separately what should happen if screening fails.
Availability note: the general guide currently names v1.106.0, but the merged backport and stable release notes confirm inclusion in v1.105.0. This guide concerns that stable release, not every feature in newer documentation.
Choose what a flagged message should do
Each check has a question, an action and a threshold between zero and one. A returned probability at or above the threshold triggers the action. The release’s configuration schema defaults to block at 0.7, so an observation pilot must explicitly choose log.
Action | When it fits | What happens when flagged |
|---|---|---|
Log | You are learning whether a policy catches the right messages. | Records the verdict and allows processing, provided no block check fires and screening succeeds. |
Block | The prohibited behavior is clearly defined and evaluation supports enforcement. | Rejects the guarded request when a block check triggers. A mistaken flag can therefore deny a legitimate request. |
These are rollout recommendations, not validated accuracy claims. The included actions do not include holding a message for a human: a review queue needs separate application work.
Logging and service availability are separate decisions
The v1.105.0 implementation defaults to fail_closed: a screening-service failure raises HTTP 502. Reading the code shows that this also applies when every check uses log. “Observation only” describes the verdict action, not a guarantee that requests cannot fail.
Choosing fail_open instead allows processing after a screening failure and records a warning; a positive blocking verdict still wins. That trades enforcement during an outage for availability. Plan separate user-facing handling for a policy rejection, normally HTTP 400, and a screening-service error. Neither choice is right for every application.
Pick a threshold from the mistakes you can tolerate
The default 0.7 is not a promise of 70% accuracy, nor a universal safe setting. A model’s probability output needs checking against labeled outcomes on the traffic where it will be used. scikit-learn’s calibration reference explains this distinction.
Build examples of violations and legitimate requests, including quoted instructions, corrections and the languages your customers use. Compare candidate thresholds on a validation set, then check separate held-out examples. Count both legitimate requests wrongly blocked and violations missed; choose based on their consequences. This follows standard threshold-selection practice, not a LiteLLM performance benchmark.
Evaluate complete requests with realistic conversation lengths. The implementation uses the highest score across screened messages and chunks for each check. A successful single-sentence demonstration therefore does not establish how often a full conversation will be stopped.
TypeSafe, one supported provider, documents limitations involving literal interpretation, arithmetic and adversarial text in Jev 1.13. Our advice is to keep exact permissions and spending limits in ordinary application controls; semantic screening should supplement them, not replace them.
Budget for screening work, not just conversations
LiteLLM sends the questions together, but screens messages separately and splits long text into overlapping segments. Tool-call text adds work too. The tagged implementation therefore does not imply one screening call per customer interaction.
For one provider example, TypeSafe’s published rate checked October 11 lists Jev 1.13.0 at US$0.042 per million input tokens, with no output-token charge. This is not a universal LiteLLM screening price.
Illustrative arithmetic, not a measured bill: one million interactions × one screening call × 1,000 billed input tokens = one billion tokens, costing $42 at that rate. Ten such calls per interaction would cost $420. Assume the 1,000-token average includes all billed input components, with no discounts, retries or additional minimum fees. Chatbot generation, hosting and human review are excluded.
The LiteLLM guide says internal guardrail calls do not create spend rows or count against the originating key/team’s spend accounting. That does not make provider inference free. Reconcile upstream usage separately, and measure added response time during a pilot.
Check what reaches the user
Streaming deserves particular care. LiteLLM’s documented behavior allows earlier answer fragments to reach the client before a later scan flags the response and ends the stream. End-of-stream screening is also not pre-delivery review. If an answer must be checked before anyone sees or acts on it, the application needs to withhold it until screening completes.
There is also an explicit coverage gap: Responses requests containing only function_call and function_call_output items, with no instructions or user text, are not screened. Do not treat this feature as complete protection for every step of an agent’s work.
Start with one policy and a bounded pilot
Confirm the installed release and access to the selected decision model. Attach the policy deliberately; it is not automatically active for every request. Use
pre_callwhen the requirement is input screening before generation. Configuration guidance.Choose log explicitly, agree on failure behavior, and inspect actual verdict records. Record false flags, missed violations, service errors, response time and upstream usage before deciding whether to enforce.
Keep the evaluated model version, question wording and threshold together. TypeSafe’s versioning guidance warns that aliases can change; pin a version where reproducibility matters and reevaluate changes before enforcement.
For background on classification versus generated answers, see our OpenAI Decisions API guide. Here, the decision is narrower: whether a particular screening policy is ready to interrupt a real chatbot conversation.
Methodology: AI-assisted reporting and analysis based on published documentation, release records and static inspection of v1.105.0 source, checked October 11, 2026. No live gateway tests or independent accuracy or latency benchmarks were performed. Cost figures above are hypothetical.
