Microsoft Launches MAI Models, Expanding Commercial Multimodal Capabilities in Image and Speech
Microsoft is continuing to expand its portfolio of in-house models through Microsoft AI, reducing enterprises’ reliance on a single external provider when adopting generative AI. The latest additions cover image generation, speech recognition and speech synthesis. Integration with Microsoft Foundry, Copilot, Bing, PowerPoint and Azure Speech allows companies to deploy visual, listening and speaking capabilities within existing workflows.
Microsoft announced on April 2, 2026, that MAI-Transcribe-1, MAI-Voice-1 and MAI-Image-2 were joining Foundry. The transcription model supports 25 languages and starts at $0.36 per hour. The voice model can generate 60 seconds of audio in under one second on a single GPU and starts at $22 per million characters. The image model ranked No. 3 in its category on Arena.ai.
All Coverage
5 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →