
Gemini 3.8 Live API: Pricing, Voices, Setup
Summary: Gemini 3.8 Live in 30 Seconds Google promoted two voice models to general availability on September 15, 2026: gemini-3.8-live and gemini-3.8-live-extended-thinking. Both are audio-to-audio. No speech-to-text model in front, no text-to-speech engine behind. Raw audio in, raw audio out. Pricing is quoted per minute as well as per token: $0.005/min for audio input, $0.018/min for audio output. There is a free tier. The only difference between the two models is thinking. The standard one rejects thinkingLevel; Extended Thinking accepts low, medium, high. Extended Thinking takes #1 on Artificial Analysis’ Speech to Speech Quality Index at 82.6. Plain 3.8 Live sits second in the Speech Agent Arena. It auto-detects and switches between 97 languages mid-conversation. The trap: without compression, audio-only sessions cap at 15 minutes and audio-plus-video sessions at 2 minutes. Google’s September did not end with 3.8 Flash. Today, September 15, 2026, two new stable models landed on the Live side of the Gemini API: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. ...