Meta Unveils Muse Voice Transcribe with Native Support for Five Indian Languages

Photo: IANS

Meta has unveiled Muse Voice Transcribe, its first real-time audio-perception model, bringing advanced speech-to-text capabilities and native support for five major Indian languages.

Developed by Meta Superintelligence Labs, Muse Voice Transcribe is designed to handle real-time transcription while also identifying and separating multiple speakers. Meta said the model can distinguish more than 20 voices in recordings lasting over an hour, while supporting multilingual conversations, code-switching and speaker diarisation through a single model without requiring a separate post-processing stage.

The model has been trained on more than 70 languages spoken across multiple countries, with 25 languages validated for support at launch. According to Meta, Muse Voice Transcribe ranked first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026.

The model is already being used for dictation in Meta AI for Mac and Muse Code and is also available through the Meta Model API. Its capabilities include real-time automatic speech recognition, speaker diarisation, endpointing and multilingual transcription.

One of its notable features is its ability to handle code-switching, a common occurrence in multilingual countries such as India where speakers frequently shift between languages within the same conversation. The system can also use language, keyword and contextual information to improve transcription accuracy.

Meta has introduced what it calls an "adaptive delay" mechanism to address one of the biggest challenges in real-time transcription: balancing speed against accuracy.

Normally, a speech-recognition system can improve accuracy by waiting longer before producing a transcription, but that increases latency. Muse Voice Transcribe instead dynamically adjusts how long it waits before predicting each word, depending on how difficult that word is to recognise.

The company said this capability was developed using reinforcement learning, with word-error-rate and delay-related rewards combined to optimise the balance between accuracy and response time.

Technically, Muse Voice Transcribe is an autoregressive multimodal model belonging to Meta's Muse Spark family. Audio is processed in 80-millisecond chunks, with each chunk converted into a single soft token. At every stage, the model determines whether it should continue listening for additional audio or generate a text token.

According to Meta, this adaptive approach allows Muse Voice Transcribe to achieve an improved speed-versus-accuracy balance, particularly in applications where users expect transcription to appear almost immediately while still maintaining high accuracy.

Native support for five Indian languages could make the technology particularly relevant to India's multilingual digital ecosystem, where users frequently communicate in regional languages or switch between English and an Indian language during conversations.

The development also highlights the growing competition among technology companies to build speech and audio AI systems capable of handling real-world conversations rather than simply converting carefully spoken audio into text.

With real-time multilingual transcription, speaker identification and code-switching built into a single model, Meta is positioning Muse Voice Transcribe as a broader speech-understanding system rather than merely another speech-to-text engine.

Follow Us
Read Reporter Post ePaper
--Advertisement--