Posts tagged “voice-agents”
Gemini 3.8 Live Avatar Is GA. The Face Has to Earn Its Place.
Google’s Gemini 3.8 Live Avatar is enterprise GA. Here’s what speaking-time pricing, short sessions and identity controls mean for agent builders.
Gemini 3.8 Flash TTS Separates the Script From the Speaker
Gemini 3.8 Flash TTS and Flash-Lite bring reusable voices to AI agents. Compare pricing, voice expiry, speech accuracy, and migration tradeoffs.
Gemini 3.8 Live Extended Thinking Makes “Done” a Three-Clock Problem
Google’s new Gemini Live models make background work first-class. Builders now need separate speech, interaction, and transaction states.
OpenAI Watermarks Supported GPT-Live Audio. The Evidence Trail Is Still Yours
OpenAI adds SynthID to supported GPT-Live audio and opens verification access. Here is the evidence workflow voice-agent builders now need.
OpenAI Presence Turns Agent Deployment Into Release Engineering
OpenAI Presence packages policies, evals, escalation, and Codex-assisted updates into a managed agent platform. The moat—and lock-in—sits above the model.
OpenAI’s Audio API Retirement Turns Aliases Into Migration Debt
OpenAI will retire legacy audio and Realtime API models in January 2027. Here is the migration map, cost trap, and production test plan.
GPT-Realtime-2.1 Makes Voice Agents a Production API Bet
OpenAI's GPT-Realtime-2.1 release gives API voice agents a sharper cost, routing, and eval model across WebRTC, WebSocket, and SIP.