Microsoft completes full end-to-end pipeline for voice agents: real-time transcription as you speak, voice generation latency as low as 45ms.
Beating AI Express News: Microsoft has rolled out three audio models in one go: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The first is designed for real-time speech-to-text, while the latter two handle text-to-speech, primarily targeting low-latency voice agents.
MAI-Transcribe-2-Streaming can output text continuously before a speaker finishes talking, supporting 60 languages and automatic language detection. Per Artificial Analysis’ streaming speech-to-text leaderboar
@tonmelonfest
Beating AI Express News: Microsoft has rolled out three audio models in one go: MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash. The first is designed for real-time speech-to-text, while the latter two handle text-to-speech, primarily targeting low-latency voice agents.
MAI-Transcribe-2-Streaming can output text continuously before a speaker finishes talking, supporting 60 languages and automatic language detection. Per Artificial Analysis’ streaming speech-to-text leaderboar
@tonmelonfest