Meta has introduced Muse Voice Transcribe, its first real-time audio model designed for live dictation and transcription across multiple speakers and languages.
Meta says the model can handle more than 20 speakers, switch between languages automatically, and understand code-switching, where people use words from different languages within the same sentence.
Trained Across More Than 70 Languages
Meta CEO Mark Zuckerberg demonstrated the model in a video showing it automatically identifying different speakers and switching between languages during a conversation.
According to Zuckerberg, Muse Voice Transcribe uses adaptive delay to improve transcription accuracy. It waits longer before committing to difficult words while processing easier words more quickly.
The model was trained across more than 70 languages, with 25 languages validated at launch.
Meta says it can also handle noisy and complicated real-world audio, mid-sentence language switching, and sessions lasting around an hour with more than 20 speakers.
Zuckerberg shared the demonstration after recently returning to X following roughly three years without posting on the platform.
Muse Voice Transcribe is MSL’s first real-time audio perception model — rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model. pic.twitter.com/LViMDSkbim
— Mark Zuckerberg (@finkd) September 1, 2026
Meta Takes on Google Gemini
Muse Voice Transcribe arrives less than a week after Google introduced Gemini 3.5 Transcribe, which offers similar audio transcription capabilities.
Google plans to integrate its model into Android and eventually Chrome.
Meta has not said whether Muse Voice Transcribe will eventually become part of its major consumer services.
Available Through Meta AI and APIs
For now, users can access the technology through Meta’s recently launched Meta AI Mac app.
Because the Mac app can provide voice features to other applications, Muse Voice Transcribe can also power dictation in other services through the app.
Developers can access the model through Muse Code and Meta’s Model API.
Meta has priced the service at $3 per 1,000 minutes of audio.
A demonstration version of Muse Voice Transcribe is also available through Meta’s research blog.
Muse Voice Transcribe is the latest release from Meta Superintelligence Lab (MSI). In recent weeks, the group has also introduced Meta’s first dedicated coding agent, an open-weight model, and the Meta AI Mac app.
The post Meta’s Muse Voice Transcribe 70+ Languages and 20+ Speakers Simultaneously appeared first on ProPakistani.
