Meta has launched Muse Voice Transcribe, its first real-time audio perception model, built with speaker diarization and endpointing in a single system.
The model handles dictation and transcription for more than 20 speakers and can seamlessly manage multiple languages at once.
Chief executive Mark Zuckerberg demonstrated the tool in a video, showing it distinguishing between speakers in real time and switching between languages within the same conversation.
Trained across dozens of languages
The model was trained across more than 70 languages, with 25 validated at launch, and can handle messy, real audio as well as hour-long sessions with more than 20 speakers.
The system also copes with code-switching, automatically picking up when speakers blend words from multiple languages mid-sentence.
Pricing and availability
The model is available now through Meta's Model API and within Muse Code, and can also be used in Meta's recently released AI Mac app.
It is priced at $3 for 1,000 audio minutes, with a demo version also available online.
Part of a wider reset
The release follows months of delays around Meta's flagship internal model and a broad reorganisation of its artificial intelligence effort into a single labs structure, changes pursued to close gaps with leading rivals.
A crowded field
The launch lands in the middle of a busy competitor cycle.
It arrived less than a week after Google introduced its own audio transcription model with similar capabilities, while Google has separately been promoting its latest fast-response model and Fei-Fei Li's World Labs has unveiled a new world model of its own.
Where the contest goes next
Meta's latest launch shifts the contest towards specialised and device-friendly systems, rather than broad general-purpose chatbots.
Its real-world impact will become clearer as benchmark results and product integrations surface over the coming weeks.