8news

Tech • AI • Robotics

VIDEO
ENFR
TodayShortsTop StoriesYour topicFor youTopicsAll videosYT channelsArchivesSearchFavorites

How to build with Gemini 3.5 Transcribe

7/10
GoogleGoogle for DevelopersAugust 26, 2026 at 09:26 PM4:49
Audio player
0:00 / 0:00

TL;DR

Google has launched its first Gemini-based transcription model, offering batch and live speech-to-text with improved accuracy for items such as email addresses, phone numbers, mixed languages and measurement units.

KEY POINTS

Launch across two APIs

The new Gemini 3.5 Transcribe model is available for standard transcription through the Interactions API and for real-time use through the Live API. The release targets applications that need low-latency speech recognition as well as post-processed final transcripts.

LLM-based transcription

Unlike conventional speech-to-text systems, the model is built as an LLM-based transcription system. That design is intended to improve recognition of difficult spoken content such as alphanumeric strings, names, and email addresses, which often cause errors in standard transcription pipelines.

Custom vocabulary support

Accuracy can be improved by providing a custom vocabulary, such as the names of people expected in a meeting. In one demonstration, a nonstandard personal name was correctly recognized after being supplied as prior context, showing how domain-specific hints can reduce mishearing.

Self-correction in final transcripts

A notable feature is the model’s ability to revise earlier output in the final transcript. When an email address was initially captured incorrectly, the system later corrected it after processing more context, producing a more accurate final result than the first live pass.

Formatting for emails and phone numbers

The model was shown recognizing spoken contact details and formatting them correctly, including email addresses and phone numbers. It identified a U.S. phone number in the expected format and also handled a Singapore number, reflecting support for region-specific conventions rather than simple digit capture.

Language hints and multilingual recognition

Users can set expected language hints to improve transcription quality, but the model can still automatically recognize speech in more than 70 languages. Even when configured for English, it remained able to detect and transcribe words from German, indicating support for multilingual and code-switched conversations.

Handling units and numeric meaning

The system also aims to interpret spoken measurements and number-related context more accurately, including units such as meters and centimeters. It is designed to reduce ambiguity in cases where language differences can affect the meaning of large numbers and measurement terms.

CONCLUSION

The launch positions Gemini 3.5 Transcribe as a more context-aware alternative to traditional speech recognition, especially for live transcription where correctness of names, numbers and multilingual speech matters most.

Explain this
Full transcript

More from Google