AI News HubLIVE
サイト内リライト1 分で読了

翻訳待ち:Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's voice-controlled AI features, without struggling with background noise or when your speech is interrupted. Gemini 3.5 Transcribe is a completely new addition to the Gemini family, and its introduction comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June. Google says that 3.5 Transcribe "repres … Read the full story at The Verge.

ソースThe Verge AI著者: Jess Weatherbed

AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。

Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google’s voice-controlled AI features, without struggling with background noise or when your speech is interrupted. Gemini 3.5 Transcribe is a completely new addition to the Gemini family, and its introduction comes as we’re still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June. Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.” Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps. Transcribe is launching alongside 3.5 Live and 3.5 Live Experimental, which build on the existing speech recognition tech that powers Gemini’s voice chat mode. Gemini 3.5 Live is better at handling mid-sentence interruptions, language recognition, and live visual processing, while Gemini 3.5 Live Experimental goes further by narrating its progress step by step in real time while it tackles reasoning on more complex tasks. These Gemini Audio updates are rolling out starting today in English for all macOS Gemini app users, and the Rambler dictation feature on Android in select countries and languages. It’s also available for developers in public preview in the Gemini API via AI Studio and Antigravity. Google says that Chrome support is coming soon.