Google Introduces Gemini 3.5 Transcribe AI Model That Cleans Up Filler Words
The newly developed Gemini 3.5 Transcribe model filters out hesitation expressions during speech, recognizes custom terminology, and can separate audio for up to three speakers.
Google has launched Gemini 3.5 Transcribe, a new AI model that automatically cleans up unnecessary pauses in audio recordings and supports more than 85 languages.
Advanced Audio Transcription Features
The Gemini 3.5 Transcribe model announced by Google shows significant progress in multilingual performance and word error rates compared to the previous Chirp 3 model. The model allows for automatically filtering out filler words in audio recordings and formatting the text naturally.
By defining a custom vocabulary in the system, users can ensure that technical terms and custom spellings are accurately transcribed without the need for manual correction. Additionally, speaker diarization for up to three speakers can be performed on pre-recorded audio, and word-level timestamps can be added.
Availability and Developer Support
Gemini 3.5 Transcribe has been made available in English on the macOS Gemini app and in specific countries and languages for the Rambler dictation feature on Android. Developers can access the model through the Gemini API public preview on AI Studio and Antigravity.
Google announced that the feature will also be coming to the Chrome browser soon. The company also continues to hold back the release of the Gemini 3.5 Pro model, which was promised in June.
Postponement Announcement for Other Models
The announcement process had included information that Gemini 3.5 Live and Gemini 3.5 Live Experimental updates would also be released. However, Google later issued an update stating that only the 3.5 Transcribe model was being offered today, with no new date provided for the other models.