Microsoft announced the launch of two self-developed AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, on July 23, both of which have entered the public preview stage. The two models are aimed at high-quality image generation and high-concurrency voice interaction scenarios. Microsoft emphasized that their training data has been cleaned and is traceable, and they do not rely on third-party model distillation. From the bottom-up design, they directly serve Microsoft product users.

Image model with highest precision, GPU cost reduced by up to 84%
MAI-Image-2.5-Pro is Microsoft's highest-precision image model, targeting high-demand scenarios such as main visual image generation, detailed editing, and text rendering within images, and supports natural language editing instructions. This model has become the default image generation model in Bing Image Creator, and is used for image-to-image editing functions in PowerPoint. Compared to GPT-Image-2, it can reduce GPU costs by up to 84%. After deployment in OneDrive, the success rate of relevant scenarios increased by 26%, P95 latency decreased by about 25%, and efficiency in medium-load production environments increased by 2.5 times.
Regarding pricing, it charges $5 per million tokens for text input and $106 per million tokens for image output.
Voice model speed doubled, serving customers like T-Mobile

MAI-Voice-2-Flash is optimized for high-frequency voice applications, with a speed increase of about 2 times compared to the previous generation and a cost reduction of 32%, suitable for fast-response scenarios such as customer service centers and voice assistants. The model has been integrated into Dynamics 365 Contact Center, serving customers like T-Mobile and EasyJet, with a maximum GPU cost reduction of 89%. Microsoft has also integrated it into Azure Voice Live service, supporting developers to build voice-to-voice interactive agents.
In addition, the MAI-Transcribe-1.5 in the same series has been applied to the medical voice solution Dragon Copilot, used by 170,000 healthcare providers, processing 28 million patient consultation records last quarter, supporting 58 languages, with transcription error rates decreasing by about 50% for most languages. Microsoft revealed that the next-generation GB200 computing cluster has already gone into operation, and the MAI series models will continue to expand.
