Meta's Muse Voice Transcribe hits 3.1% live error rate
The system slices speech into 80-millisecond chunks, separates more than 20 speakers and returns a final line after 0.16 seconds. Weights stay closed.
Symbolic image: a tabletop microphone array glows on a conference table while two participants, seen from behind, lean toward it.
Muse Voice Transcribe turns speech into text while it is still being spoken, working in 80-millisecond chunks at a 3.1 percent word error rate and telling more than 20 speakers apart.
At a glance
- Muse Voice Transcribe processes audio in 80-millisecond chunks; a sentence reaches its final form after about 0.16 seconds.
- Word error rate of 3.1% in streaming mode (English), ahead of ElevenLabs Scribe v2 at 3.6% and AssemblyAI Universal-3.5 at 4.0%.
- Diarization error rate averages 17.5% across AMI-IHM, AMI-SDM and VoxConverse; compared systems land between 21.1% and 28.6%.
- Trained on more than 70 languages, 25 verified in depth at launch; keeps more than 20 speakers apart in one recording.
- Pricing is $3 per 1,000 audio minutes ($0.18 per hour), via the Meta Model API, Meta AI and Muse Code.
Muse Voice Transcribe, released by Meta Superintelligence Labs on September 1, 2026, writes speech down while it is still being spoken. The model cuts incoming audio into 80-millisecond chunks, keeps more than 20 speakers apart, and reports a 3.1 percent word error rate on English.
Why 80 milliseconds matters
An assistant that answers the moment you stop talking needs a transcript that is already finished. Muse Voice Transcribe works the audio stream in 80-millisecond slices and varies how long it holds each word before committing: easy words pass through quickly, ambiguous ones get more time. A passage settles into its final form after roughly 0.16 seconds. Meta places the model in its Muse Spark family and describes it as an autoregressive multimodal system.
On the Artificial Analysis streaming leaderboard, that 3.1 percent English word error rate ranks first. ElevenLabs Scribe v2 records 3.6 percent and AssemblyAI Universal-3.5 records 4.0 percent, though Scribe v2 commits faster at 0.14 seconds.
Telling voices apart for an hour
Continuous listening falls apart if the system cannot say who spoke. Meta reports a 17.5 percent average diarization error rate across AMI-IHM, AMI-SDM and VoxConverse, against 21.1 to 28.6 percent for the systems it compared against. The model is meant to hold more than 20 voices apart in a single recording, including sessions running past an hour, without a separate post-processing pass.
It also marks sentence boundaries and follows switches between languages inside a single sentence. Training covered more than 70 languages, of which Meta verified 25 in depth for launch. No published accuracy figures exist for the rest.
Glasses, price, and what stays hidden
The Decoder reads the release as infrastructure for Meta's pitch of a personal superintelligence: assistants in AI glasses that listen the whole time rather than waking on a keyword. That is also where the objections begin. Meta's camera glasses have already drawn privacy scrutiny in Germany, and a microphone that never stops is a harder case than a lens someone can see pointed at them.
Access runs through the Meta Model API, Meta AI and Muse Code, priced at $3 per 1,000 audio minutes, or $0.18 per hour of audio. Meta has not released the weights and has not disclosed the parameter count or the sources of its training data. Every accuracy figure above comes from Meta's own announcement and the Artificial Analysis ranking it cites; no independent replication has been published so far.
FAQ
What is Muse Voice Transcribe?
A transcription model from Meta Superintelligence Labs, released September 1, 2026, that processes live speech in 80-millisecond chunks, detects sentence boundaries and separates more than 20 speakers.
How accurate is Muse Voice Transcribe?
Meta reports a 3.1 percent word error rate on English streaming audio in the Artificial Analysis ranking, where ElevenLabs Scribe v2 sits at 3.6 percent and AssemblyAI Universal-3.5 at 4.0 percent. For speaker separation, Meta reports a 17.5 percent average error rate.
Can I download or self-host Muse Voice Transcribe?
No. Meta has not released the weights, so the model is only reachable through the Meta Model API, Meta AI and Muse Code, at $3 per 1,000 audio minutes or $0.18 per hour.