Vercel released AI SDK 7.0.135 on October 8, 2026, adding a built-in WebSocket chat transport for JavaScript and TypeScript teams building streaming chat interfaces. It lets multiple chat operations share a persistent UI-to-server connection, with response correlation, cancellation frames and stream-replay support. The server protocol and retained response storage remain application work.
Our recommendation: keep HTTP for straightforward chat; evaluate the new transport when frequent browser-side tool follow-ups or multiple concurrent chat streams justify operating a shared socket. This release supplies useful protocol machinery, not evidence that every chat will become faster or cheaper.
Which chat applications should switch?
The transport guide keeps HTTP POST as the default. WebSockets are opt-in. The following choices are architectural recommendations based on the documented behavior, not benchmark results.
Your requirement | Suggested choice | Reason |
|---|---|---|
Ordinary prompt-and-response chat | Keep HTTP initially | A persistent socket adds server and lifecycle responsibilities without an established benefit for this workload. |
Frequent client-tool follow-ups or several chat streams | Pilot WebSocket transport | Turns can reuse one socket, and request IDs separate their response streams. |
Recover a response after a page reload | Evaluate HTTP resume first | WebSockets are not required for retained-stream recovery. Both designs need application-owned storage. |
Keep an agent running through a server restart | Design execution durability separately | A persistent connection and a retained output log do not themselves preserve running model or tool work. |
The existing HTTP resume guide already supports page-reload recovery using stored streams and POST/GET endpoints. Also, HTTP can reuse underlying connections: one request per turn does not mean one new TCP/TLS handshake per turn. Do not infer a latency saving from request counts alone.
This is not a provider realtime or audio session. It carries AI SDK UI messages between your interface and backend. Client-tool follow-ups still start another turn after the preceding UI stream finishes; they do not inject tool output into an active generation. API scope and turn semantics.
Wire the client to a protocol-aware server
1. Pin compatible packages and assign a transport owner
The registry confirms ai 7.0.135 and @ai-sdk/react 4.0.138, which depends on that AI SDK version. Both declare Node.js 22 or newer. Check the declared Zod and React peer ranges before changing an existing lockfile; the React range does not include every React 19 patch.
For an already compatible React project, this pins the two packages. It is an installation example, not a tested migration:
npm install --save-exact ai@7.0.135 @ai-sdk/react@4.0.138Import WebSocketChatTransport from ai, create a stable instance with new WebSocketChatTransport({ url: '/api/chat' }), and pass it to useChat({ transport }). The endpoint must accept WebSocket upgrades. The lifecycle contract makes the creator responsible for calling transport.close() after the final consumer is gone; hooks do not automatically close caller-supplied transports. One chat unmounting must not close a socket still used by others.
2. Implement the server contract, not just the upgrade
An existing POST/SSE handler is insufficient. The protocol reference defines client send, resume and abort frames, and server start, chunk, end, error and no-active frames. Use its exported types and validation helper, then implement these application responsibilities:
Authenticate the connection, validate its origin and authorize the requested chat against that identity. Use WSS in production. The transport’s
headersoption carries per-frame metadata, not browser handshake headers; migrating an HTTP authorization header alone does not authenticate the upgrade.Parse untrusted JSON and call
safeValidateWebSocketChatTransportRequestbefore handling it. Enforce application permissions, frame-size and connection-rate limits. A validrequestIdcorrelates an operation; it does not grant access to a conversation.Run model work server-side and wrap each valid UI chunk with the matching request ID and a contiguous sequence beginning at zero. Bind abort requests to server-owned cancellation, and retain output under the authorized chat and generation identity.
Resume the response without rerunning the generation
Recovery requires a complete replay from sequence zero, including the original assistant message ID; reconnecting is not a request to run the model again.
The resume contract omits lastSequence because Chat reconstructs message state with a fresh parser. It needs earlier text, reasoning and tool chunks, not just the unseen tail. Preserve the original UI start chunk’s messageId so the replay replaces the partial assistant message. That UI chunk is distinct from the protocol’s start acknowledgement.
Retain the complete response log and stable message identity, with explicit expiry and access rules. Store it outside process memory when reconnects can reach another server.
On an authorized resume, replay from zero and join the live stream without an ordering gap. If no retained response exists, return
no-activeand make that state visible instead of silently starting fresh model work.Keep action deduplication separate. Replaying output should not re-execute a payment, message or other external tool action. Decide how to reconcile unfinished work before offering a retry.
That last distinction is an application-design recommendation: the tagged client implementation deduplicates chunk sequences, not external side effects. For the broader separation between a chat interface and durable execution, see AI Chatbot, Workflow or Agent? Choose by the Task.
Test the shared socket’s failure behavior
In the 7.0.135 implementation, a valid error frame fails its correlated request, while protocol failures such as an invalid chunk or sequence gap terminate the connection and fail active streams. A dropped connection also errors active streams. Later operations can open a socket, but interrupted sends are not automatically resubmitted. Our architectural inference: choose which chats share a connection deliberately, because they also share a connection-failure boundary.
Cancellation needs separate treatment. The client attempts an abort frame only on an open socket, and transmission is best-effort. Local consumption ending is not proof that server work stopped, nor does it reverse a completed tool action. Verify the server’s cancellation path, including when outbound data is queued. Abort implementation.
The outbound queue waits for the socket buffer to drain before another frame. The buffer helper uses a 1 MiB threshold and 20 ms polling by default; without bufferedAmount, that wait is skipped. The threshold is not a frame-size ceiling: a single large frame can exceed it. It also does not bound incoming data or retained response logs. Set those limits independently.
Account for hosting lifetime and memory charges
As checked on October 8, Vercel Functions WebSocket hosting is beta on all plans and requires Fluid compute. Next.js uses the experimental upgrade API documented there. Connections end at the function’s duration limit, and reconnects may land on another instance. Retained state therefore needs cross-instance access; keeping chunks does not by itself keep generation alive after the executing process ends.
For Fluid compute’s Node.js, Bun and Python runtimes, the duration limits list a 300-second default. Hobby tops out at 300 seconds; Pro and Enterprise have an 800-second generally available maximum. The extended 1,800-second beta requires supported runtime versions and function-level configuration, and excludes Secure Compute and Static IPs above 800 seconds. Check the deployed function rather than assuming a 30-minute connection.
Budget for more than active CPU. Vercel’s pricing documentation says provisioned-memory billing continues while requests remain in flight, including I/O waits. Include transfer, retained storage and model usage in the comparison; an idle socket is not necessarily a zero-cost socket.
A focused pilot before changing the default
Use the same model, prompts, tool stubs, region and concurrency for an HTTP baseline and WebSocket candidate. The following is a proposed validation plan; no results are claimed here:
Run successive turns and simultaneous independent chats. Unmount one consumer and check that other consumers continue.
Disconnect mid-text and mid-tool output; resume through another process. Check complete reconstruction, stable assistant identity and no repeated external-action attempts.
Exercise duplicate and missing sequences, invalid frames, expired authorization, retention expiry and cancellation under buffer pressure. Record the expected scope of each failure.
Compare completed-turn latency distributions, recovery success, queued bytes, memory and total cost. Define workload-specific acceptance criteria before deciding whether to migrate.
Adopt the transport when that pilot demonstrates a useful improvement or removes custom protocol work you otherwise need. Keep the HTTP path available during rollout, but do not turn a failed stream into an automatic second generation.
Methodology: AI-assisted reporting and architectural analysis based on Vercel’s release, documentation, npm metadata and version-pinned source, checked October 8, 2026. No installation, benchmark or deployment was performed. Performance gains and production reliability remain unmeasured here.
