Meta Launches Muse Voice Transcribe for Real-Time Speech Recognition
Meta Superintelligence Labs has introduced Muse Voice Transcribe, a real-time audio perception model that combines streaming automatic speech recognition, speaker diarization and endpoint detection in a single system. The unified design is intended to reduce the complexity and latency involved in linking separate speech tools, making it relevant for voice assistants, live captions and meeting transcription. The release also underscores Meta’s push to build multilingual audio infrastructure alongside its broader artificial-intelligence efforts.
Muse Voice Transcribe supports more than 70 languages and can distinguish over 20 speakers in multi-person audio, according to the release. Meta has made the model available through the Meta Model API and on Mac computers. On macOS, users can hold the Fn key to activate system-wide voice input, extending real-time transcription across applications and text fields without requiring each app to integrate a separate speech-recognition service.
All Coverage
3 original reportsThe Backstory
The history behind this eventNo historical echoes for this signal
Subscribe to Mark Radar Weekly
Every Friday, the week's strongest signals in your inbox. Unsubscribe anytime.
If you search news on Google, you can set Mark Radar as a preferred source—our coverage will show up more often in your results. Set as preferred source on Google →