Meta Superintelligence Labs on Tuesday introduced Muse Voice Transcribe as its first real-time audio perception model. Unlike a basic transcription system that processes a recording after the fact, Muse produces text continuously while identifying speakers and detecting when speech begins and ends.
The model supports audio with more than 20 speakers and recognizes speakers switching languages during a conversation. Meta says Muse was trained across more than 70 languages, with 25 extensively validated for the initial release.
Those validated languages include Hindi, Tamil, Telugu, Kannada and Malayalam, making the model potentially useful in multilingual markets such as India, where conversations may move between English and regional languages.
Muse also supports language, keyword and context biasing to help it recognize terms based on additional information available to the model.
One of Muse’s key technical features is its approach to latency. Meta says the model processes audio in 80-millisecond chunks and decides how long it needs to listen before committing to each word. Easy words can be transcribed quickly, while difficult ones get more audio context before the model makes a decision.
That timing is controlled through what Meta calls “adaptive delay,” trained using reinforcement learning. The goal is to avoid forcing the entire transcription system into a single compromise between speed and accuracy.
Meta says Muse reached the Pareto front for speed and accuracy when measured by time to final transcription. It also claims the model ranked first on Artificial Analysis’ streaming speech-to-text leaderboard and public diarization benchmarks as of Sept. 1.
What's hot at TechRepublic
- Blackpoint Cyber vs. Arctic Wolf: Which MDR Solution is Right for You?
- Why AWS Sellers Choose Deepgram Over Other Voice AI Tools
- SS&C Intralinks DealCentre AI vs. Datasite: Which platform is built for the future of dealmaking?
- SS&C Intralinks FundCentre AI vs. Juniper Square: Which platform better supports modern private markets fund managers?
- Verito vs. Rightworks: Which IT Provider Is Best for Your Firm?
Already powering Mac dictation
For consumers, the most immediate use is in Meta AI for Mac, where Muse Voice Transcribe now powers dictation features. The model is also being used in Muse Code. Developers can access it through Meta’s Model API at $3 per 1,000 audio minutes, which Meta says works out to roughly 18 cents per hour.
The major opportunity may be outside Meta’s own apps. A single model that combines transcription, speaker labeling and endpoint detection could reduce the need for developers to stitch together separate speech-processing systems.
What it could mean for voice AI
The biggest practical benefit may be that Muse is not limited to turning a clean recording into text. Its combination of live transcription, speaker labeling and multilingual recognition makes it better suited to meetings, dictation, coding and voice-driven applications.
There are still reasons to be cautious. Meta’s 70-plus language figure refers to training coverage, while only 25 languages were extensively validated at launch. Performance may therefore vary between languages and real-world recordings. Developers should test Muse with their own accents, background noise, terminology and language combinations rather than assuming equal accuracy across all 70-plus training languages.
Read more: Plaud One uses AI earbuds to record, transcribe, summarize, and act on workplace conversations without relying on a nearby smartphone.