Article

Pydantic AI 2.54.0: Streaming Concurrency Fix and Upgrade Checklist

Check which Pydantic AI streaming apps need the concurrency fix, why 2.53.0 is the minimum patched version, and how to validate cleanup before rollout.

Editorial illustration for Pydantic AI 2.54.0: Streaming Concurrency Fix and Upgrade Checklist: a document represents the research briefing. Not documentary evidence.

Pydantic AI disclosed a streaming concurrency-limit flaw on October 2, 2026. Its security advisory covers pydantic-ai and pydantic-ai-slim versions >=2.10.0 and <2.53.0 when streaming through ConcurrencyLimitedModel or limit_model_concurrency. A long-lived shared limiter can lose capacity until later requests cannot proceed.

2.53.0 is the first fixed version. Version 2.54.0, published on October 3 UTC, is the latest stable release checked for this guide and is available for both pydantic-ai and pydantic-ai-slim. If you already run 2.53.0, you have this particular patch; 2.54.0 adds other changes to assess.

Which deployments need prompt action?

Use the advisory’s affected range and the 2.53.0 security notes to separate exposed streaming paths from unrelated uses of Pydantic.

Deployment or request path

Decision for this defect

Affected version; streaming through a model-level limiter reused across requests

Prioritize patching, especially on network-accessible endpoints.

Affected version; stream_text() through a model-level limiter finishes normally

Still affected with default debouncing. Completion alone is not a safety check.

Only agent-level max_concurrency, without a model-level limiter

Outside this defect.

Exclusively non-streaming model requests

Outside this defect.

Pydantic AI v1

The release notes say v1 is unaffected.

Already on 2.53.0

This fix is present; evaluate 2.54.0 separately.

The maintainers rate the issue High, CVSS 3.1 7.5, for availability impact. Their advisory does not establish exploitation in the wild or an affected-deployment count. This is not a blanket advisory for the separate pydantic validation package.

Why a completed response can leave the pool blocked

The patch explanation describes a task-ownership mismatch: a stream acquires a slot on one task, but cleanup can run on another. The previous limiter rejects that release. The correction uses a semaphore so release need not happen on the acquiring task.

Client disconnects are one trigger, not the whole problem. The release’s security note also covers early iteration exit, consumer exceptions, cancellation and full stream_text() consumption with default debouncing. Disabling debouncing is therefore not a comprehensive remediation.

Deployment implication: increasing the concurrency limit cannot repair missing releases. A response-content assertion can pass while capacity remains occupied. Completed responses are not enough: stream cleanup must return shared concurrency slots so later requests can proceed.

Upgrade the deployed environment, then check sharing

Start with the interpreter used by the service and the application source tree. These inventory commands are suggestions, not checks performed against your deployment:

python -m pip show pydantic-ai pydantic-ai-slim
rg -n "ConcurrencyLimitedModel|limit_model_concurrency|ConcurrencyLimiter|max_concurrency|run_stream|stream_text" .
  1. Trace model factories and fallback wrappers as well as direct constructors. Record which limiter objects persist across requests and which agents or models share them; matching numeric limits do not imply a shared object.

  2. Pin 2.54.0 for the distribution you actually use through your existing dependency and lockfile workflow. Keep provider extras on slim installations; do not replace slim with the full package by accident. PyPI’s 2.54.0 metadata requires Python 3.10 or newer.

  3. Rebuild and redeploy the workers. Check resolved package versions in the resulting runtime, then run python -m pip check. A changed lockfile does not establish that every serving process uses the new code.

For a plain pip environment using the full package, the pinned upgrade is:

python -m pip install --upgrade "pydantic-ai==2.54.0"
python -m pip check

The fix also changes supported limiter arrangements. The release-pinned wrapper implementation rejects the same limiter instance at the requesting agent and its model wrapper, or at nested model wrappers, with UserError. Use separate instances or limit at one layer. Raising capacity does not remove that identity check.

Sharing a pool among sibling models remains supported by the model documentation. For manual limiter use, match each successful acquire() with one release(); repeated acquisitions on one task each consume a slot. Custom limiters must support cross-task release, as the patch migration notes explain.

Proposed regression checks before rollout

Adapt the upstream streaming regression to staging. Its deterministic FunctionModel case uses a one-slot limiter, checks that occupancy returns to zero, and sends another request through that same limiter. The test is marked for asyncio; it is not proof of every transport or backend.

  • Record package versions, async backend, limiter configuration and a declared timeout. In an isolated test, cover full default-debounced text consumption, an early iteration break, a consumer exception and cancellation.

  • After the stream context exits and cleanup completes, require running_count == 0 with no other work in flight. Then require another request using the same limiter to finish within the timeout. Check both assertions after each ending.

  • Exercise an early client disconnect through your actual streaming endpoint in staging. Include queued-request cancellation and any custom limiter implementation; a library-only fixture does not cover your server’s request lifecycle.

  • Keep inputs, outputs, exceptions, timestamps and environment details with the test result. These are proposed acceptance checks, not tests performed for this article.

For a 2.54.0 rollout, also inspect the compatibility and bug-fix notes for features you use. Unexpected BackgroundTools exceptions now end the run; streaming usage includes tools executed after the final result; HTTP-client pool defaults also change. These are separate from the slot-release patch.

For broader release planning, the related agent SDK runtime-migration guide discusses dependency and runtime checks for a different SDK. It is context, not evidence for this Pydantic defect.

If the upgrade is temporarily blocked

The maintainer-supported workarounds are to use agent-level max_concurrency instead of the model wrapper, or avoid streaming through that wrapper. Merely adding an outer agent limit leaves the affected inner path in place.

There is a sizing trade-off: agent-level limiting counts whole runs, while model-level limiting counts model requests. A run can include tools and multiple model calls. When moving the control, preserve any intended aggregate cap across agents; independent per-agent limits are not a shared pool.

During canary rollout, compare active work with limiter counters such as running_count, waiting_count and available_count, alongside cleanup exceptions and queue latency. Occupied slots after all corresponding work has ended deserve investigation even when the last response looked correct.

Methodology: Prepared with AI assistance from the maintainer advisory, release notes, package metadata, documentation and release-pinned source and test definitions. No hands-on tests were run for this guide, and no production outcomes are claimed. Availability was rechecked on October 3, 2026. Release dates use GitHub’s UTC publication metadata, not the earlier dates embedded in release titles.