Recently, Google released its latest technological achievements in global language intelligence, announcing that its AI technologies and products can now support more than 300 languages, covering 7.1 billion people and 86% of the population. Google stated that a true language large model should not only complete simple text translation but also deeply understand the ways humans communicate in real life, which are full of tone, emotion, and cultural nuances.
In terms of technical architecture, traditional speech recognition usually adopts a multi-step rigid pipeline that converts audio to text, processes the text, and then converts it back to audio. This approach strips away the crucial intonation, rhythm, emotion, and context present in human conversations. To restore the authentic quality of communication, Google has driven a transition toward native audio intelligence, allowing models like Gemini to process and understand audio directly.

Gemini 3.5 Live Translate achieves real-time spoken translation across 70 languages and over 2,000 language pairs, naturally capturing code-switching and emotional cues; Gemini 3.5 Transcribe is a high-precision speech-to-text model that outputs well-formatted text even in noisy environments or with complex terminology. It also powers the Rambler feature on the Android keyboard Gboard, supporting the removal of verbal filler words, grammar correction, and direct rewriting with voice commands, as well as seamless switching between languages.
To extend AI capabilities to more low-resource languages around the world, Google officially launched the "1,000 Languages Initiative," aiming to support the top 1,000 most widely used languages globally. In this context, the Universal Speech Model (USM), trained on 12 million hours of audio, uses cross-language transfer learning technology to apply patterns learned from high-resource languages to languages with scarce training data. These research efforts build upon 25 years of open research and over 400 peer-reviewed speech papers.
In terms of language data collection, due to the natural bias of internet content towards a few dominant languages, Google abandoned the traditional single-source crawling model and instead obtained authentic cultural contexts through local open data collaborations.
This includes the WAXAL large-scale open-source speech dataset, which covers 27 languages of sub-Saharan Africa, created in collaboration with multiple organizations; the "Project Vaani," which collected over 30,000 hours of speech in 109 languages using a regional anchoring approach, developed in partnership with the Indian Institute of Science and Bhashini; and the "Amplify Initiative," which gathered multimodal data points from over 1,600 local experts and 20 universities across four continents around the world.
In addition, Google introduced an interactive language exploration tool called Language Explorer, used to continuously visualize and map the LinguaMeta data warehouse of over 7,000 languages worldwide. These technological achievements are empowering global farmers, healthcare workers, and teachers through public welfare projects such as Digital Language Inclusion Centers.
Facing the reality that over 3 billion people still lack stable internet access, Google developed the lightweight open-source translation model family, TranslateGemma. This series of models runs efficiently on edge devices for 55 languages, achieving high-quality translation without connecting to the cloud. Additionally, for populations in low-resource areas who widely use feature phones, Google piloted the "Ask Viamo Anything" (AVA) voice AI assistant in Rwanda in collaboration with organizations like Viamo, successfully answering over 2 million questions for feature phone users using the Gemini model.
In terms of accessibility design and cultural landmark pronunciation details, traditional tools often force non-standard speakers to adapt to technology. Google designed accessible interactions from scratch, introducing Sign Language to Text (SL2T) technology, trained on over 50 sign languages and supporting American Sign Language to English conversion, which powers real-time transcription on Gboard and Pixel 11. Meanwhile, Google deepened its collaboration with Māori language experts in New Zealand to optimize localized pronunciation of place names and streets in maps, embedding authentic cultural contexts and accurate pronunciations into text-to-speech models.
After 20 years of AI language development, Google's language technologies have been deeply integrated into nine core platforms, including Search, Android, Chrome, YouTube, and Google Play. Google emphasized that the core significance of technology is not to reduce the diversity of human expression, but to expand the boundaries of expression, helping more people participate, express themselves, and be understood by the world while respecting local cultures and real-world contexts.
Join Now