On September 20th, during the research achievements exhibition session of the conference on the high-quality development of teacher education and the 70th anniversary of the establishment of Qinghai Normal University, the first Tibetan multi-language full-modal AI input method in China - Zhida AI Input Method and its corresponding Tibetan font set were officially launched.

The input method was developed by the National Key Laboratory of Tibetan Intelligent Technology. It relies on the laboratory's proprietary large model and related AI infrastructure, supporting all terminals such as mobile phones and computers. It supports multiple modalities including text, voice, and images, and can synchronize the user's dictionary and input habits through a unified account.

Three regional dialects' speech and Tibetan-Chinese-English mixed OCR

Professor Duola from Qinghai Normal University and Executive Deputy Director of the National Key Laboratory of Tibetan Intelligent Technology said that the Zhida AI input method is developed based on the core technology of the Zhida large model. In terms of input methods, its text input strictly follows the national standard Tibetan keyboard layout and supports Tibetan Latin transcription to lower the learning threshold. Voice input covers the written and spoken language of the three major dialects of U-Tsang, Amdo, and Kham, and also supports mixed Chinese, Tibetan, and English voice input. OCR recognition input can take photos or screenshots to recognize mixed Chinese, Tibetan, and English texts, suitable for scenarios such as digitalization of documents, organization of teaching materials, and office work.

The input method's dictionary covers terminology from 23 disciplines including education, law, and literature, and can automatically remember users' input habits, commonly used words, and expressions. On the same day, the development team also publicly released a Tibetan font set compatible with the input method, including four styles: bold, white, artistic, and brush script.

The launch event also showcased the Zhida large model: this model has launched nine versions, with parameter sizes ranging from 800 million to 100 billion. It deploys three main models online, supports translation between 56 languages, as well as image and video understanding, and can be applied to fields such as education, research, and government affairs.