

Meta introduced Muse Voice Transcribe, a new AI model designed to listen, transcribe, and distinguish between speakers in real time. The model supports streaming transcription, can identify more than 20 speakers in a conversation, and can detect when a person finishes speaking while audio is being captured.
Meta says Muse Voice Transcribe can also switch between languages during a conversation. The model can be fine-tuned to recognize specific names, places, and other terms, adding another layer of flexibility for transcription tasks.
Meta stated that Muse Voice Transcribe processes incoming audio in 80-millisecond slices. It uses an adaptive delay system to determine when enough information is available to transcribe a word. Easier segments are processed almost instantly, while more complex phrases receive slightly more processing time.
The company trained the system using reinforcement learning, balancing transcription accuracy with lower latency. Meta claims this approach helps improve the trade-off between speed and accuracy.
The same architecture also handles speaker changes and detects conversational endpoints. This allows transcription and speaker tracking to operate within a single system.
Also Read: Muse Glimmer: Complete Guide to Meta’s Open Agentic AI Model
According to Meta, Muse Voice Transcribe was trained on over 70 languages, with 25 languages undergoing extensive validation. These include Chinese, French, Hindi, Japanese, Spanish, and Vietnamese.
Meta also demonstrated the model switching between English and Mandarin within the same sentence. The company showcased its ability to process long recordings involving multiple speakers without requiring manual cleanup.
Meta sees Muse Voice Transcribe as a building block for more personalized AI assistants that can understand accents, interruptions, overlapping speech, and multilingual conversations.
The model is rolling out immediately across several Meta products, including voice dictation in Meta AI and the company’s Muse Code tool. It is also available through Meta’s Model API and Meta AI for Mac.
Meta’s launch comes as real-time speech systems increasingly move beyond basic transcription, with speaker identification, multilingual conversations, and faster response times becoming key parts of AI-powered voice tools.