Tool: Gemini Live audio
Simon Willison’s Weblog
Subscribe
15th September 2026
Tool
Gemini Live audio
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking today - two new speech-to-speech models that are a similar shape to OpenAI's GPT-Live family.
I pointed GPT-6 Astra Extra High at the documentation and had it build me this web UI for trying out the new models. You can select a model and voice preset, enter an optional system prompt and then start a voice conversation through your browser, including the ability to interrupt the model while it is talking.
The implementation uses no libraries. It connects to the wss://generativelanguage.googleapis.com/ws/google.ai.generativelanguage.v1alpha.GenerativeService.BidiGenerateContent?key=... WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.
Here's the Gemini Live tutorial for getting started with that WebSockets API.
Recent articles
Generating running routes with GPT-6 Astra and ChatGPT Work - 12th September 2026
OpenAI agents attacked RubyGems back in May - 12th September 2026
Some thoughts on the Navier–Stokes Millennium Prize Problem - 8th September 2026
This is a beat by Simon Willison, posted on 15th September 2026.
google 416
tools 78
websockets 21
generative-ai 1,981
llms 1,947
gemini 196
llm-release 231
speech-to-text 21
Monthly briefing
Sponsor me for $10/month and get a curated email digest of the month's most important LLM developments.
Pay me to send you less!
Sponsor & subscribe
Disclosures
Colophon
©
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026