On October 1, 2026, Microsoft announced MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. For voice-agent builders, the release adds streaming speech recognition and multilingual speech generation, with separate prices for listening and speaking.
The first deployment decision is preview suitability. Microsoft’s transcription documentation and voice documentation label these Azure integrations public previews, without an SLA and not recommended for production workloads. Treat this as an evaluation path, not an automatic replacement for a production call stack.
What each model does—and what it costs
MAI-Transcribe-2-Streaming turns incoming audio into text; MAI-Voice-2.1 and Flash turn response text into speech. These are separate stages of speech recognition and speech generation. Your application still supplies the reasoning model, tool execution and conversation state.
Model | Role and published language coverage | Listed price (USD) |
|---|---|---|
MAI-Transcribe-2-Streaming | Streaming speech-to-text; 60 languages | |
MAI-Voice-2.1 | Text-to-speech; 23 languages | |
MAI-Voice-2.1-Flash | Text-to-speech; 23 languages |
Language counts come from Microsoft’s release; prices match Vercel’s model listings checked October 1. Microsoft says the transcription rate is introductory through the end of 2026. A replacement rate for January 2027 was not established in the sources reviewed.
Microsoft positions Flash for interactive speech and standard Voice 2.1 for expressive long-form output. That makes Flash a reasonable first evaluation candidate for short agent replies, while narration warrants a comparison with standard. This is a selection heuristic, not a listening-test result.
Choose an access route before copying an example
Direct Azure WebSocket: the Realtime guide describes a Foundry resource and model deployment using an OpenAI Realtime-compatible protocol. Its MAI-specific event behavior matters; compatibility does not make every Realtime feature interchangeable.
Managed Azure client: the Speech SDK guide specifies version 1.52.0 and provides connection management and recovery. Choose this route if you want the SDK to own those mechanics.
Vercel AI Gateway: the October 1 changelog lists all three models. The speech-to-text and text-to-speech guides still describe gradual beta access, so check your team’s catalog before committing to an integration.
Check regions for the chosen route. The Realtime region table lists South India, while the SDK table lists Southeast Asia instead; both list Sweden Central and Central US. Do not merge those lists into a universal availability promise. Account eligibility, preview-specific concurrency and regional capacity were not verified through an authenticated request.
Keep provisional text separate from confirmed text
For a direct WebSocket implementation, follow Microsoft’s session and event contract:
Configure before sending audio. Use raw mono, little-endian PCM16 at 16 or 24 kHz, without WAV headers. Settings cannot change after audio starts; sessions have a one-hour ceiling.
Keep two text buffers.
conversation.item.input_audio_transcription.deltaadds finalized text verbatim.conversation.item.input_audio_transcription.intermediatereplaces the provisional suffix relative to the latest delta; it is not another string to append permanently.Choose the commit boundary in your client. Server-side turn detection and automatic commit are not supported on this route.
input_audio_buffer.committedacknowledges the commit;conversation.item.input_audio_transcription.completedsupplies the final transcript for that interval.
Application recommendation: use tentative words for discardable work, such as a speculative lookup. Before changing an order or sending a message, resolve the caller’s final parameters and apply your authorization rules. A finalized transcript is not proof that an external action succeeded—a distinction also covered in our guide to voice-agent interaction and action lifecycles.
The SDK has a different interface: recognizing exposes intermediate text and recognized returns segment results. Its recovery guidance says unconfirmed audio may be replayed; do not add a resend-all loop while recovery is underway. Close the input stream and let pending results drain. Never promote partial text to a final result merely because the connection failed.
Match the spoken reply to a supported voice
Recognition coverage does not establish reply coverage: the 60-language recognizer is broader than the 23-language voice models. Build a language-to-voice mapping and an explicit fallback, such as text or handoff, for unsupported spoken replies.
Start with a prebuilt voice. Azure’s voice guide uses a model-suffixed SSML name such as en-US-Harper:MAI-Voice-2.1-Flash; Vercel’s TTS guide selects the model through gateway.speechModel() and documents locale-bearing names such as en-US-Harper. Keep each route’s configuration intact. Cloning is a separate, gated workflow requiring consent, not a prerequisite for this integration.
Also distinguish automatic language recognition from metadata your app receives. The streaming Speech SDK results do not expose detected-language labels, confidence scores or word timestamps. Do not design language routing around a field that this interface does not return.
A worked speech budget
The calculation below uses the listed rates above. Let H be billable input-audio hours and C be billable synthesis characters. These are independent workload inputs—not a conversion from call duration to generated text.
Flash speech subtotal (USD) = 0.54 × H + 15 × C / 1,000,000
Standard speech subtotal (USD) = 0.54 × H + 22 × C / 1,000,000Assume 1,000 billable audio hours and 10 million synthesized characters:
Cost component | With Voice 2.1 | With Voice 2.1 Flash |
|---|---|---|
Recognition: 1,000 × $0.54 | $540 | $540 |
Synthesis: 10 million characters | $220 | $150 |
Speech subtotal | $760 | $690 |
Calculated saving: Flash lowers the synthesis line by (22 − 15) / 22 = 31.8%. In this workload, it lowers the speech subtotal by $70 / $760 = 9.2%. The recognition bill is unchanged; the amount of generated text determines how much the substitution saves.
This is an illustrative speech-only budget, not a full agent quote or a managed Voice Live quote. It excludes reasoning-model usage, tools, telephony, hosting, additional retry usage, custom-voice charges and taxes. Confirm your route’s billing rules for silence, resubmissions and character accounting before forecasting an invoice.
What to measure in a pilot
Microsoft says the recognizer produces its first hypotheses just over 100 milliseconds after receiving audio. That is a vendor claim about provisional transcription, not a measured caller-to-reply time. Endpointing, finalization, reasoning, tools, synthesis, network delivery and playback still contribute to the experience.
A useful evaluation plan would:
Use consented, versioned recordings with names, accents, background noise, corrections and language switches; score final transcripts separately from provisional display behavior.
Measure from the caller’s end of speech to audible reply, while recording each intermediate stage. Exercise reconnects and session rollover rather than judging only an uninterrupted demo.
Record model, route, SDK version, billable audio and generated characters. Compare actual charges and response quality before choosing standard or Flash.
The practical starting point is one enabled preview route, one prebuilt voice and explicit transcript state handling. Validate those pieces together before expanding language coverage or connecting write-capable tools.
Methodology: AI-assisted reporting and analysis based on Microsoft and Vercel announcements, documentation and pricing pages checked October 1, 2026. The cost example is arithmetic on stated assumptions. No audio inference, SDK execution, latency benchmark or listening test was performed.
