PolyAI Launches Dialog-RSN-1, an Audio-Native Voice Model With Sub-300ms Latency
PolyAI's new call center voice model reasons directly over raw audio instead of a transcribe-then-generate pipeline, cutting median response time to 280 milliseconds, about a third of OpenAI's GPT Realtime-2.1.
PolyAI has released Dialog-RSN-1, a voice dialog model built for call centers that reasons directly over raw call audio rather than routing it through a separate speech-recognition step before generation. The model keeps text-to-speech as a separate module so customers can still control voice, accent, and tone.
- Median response latency of 280ms, 500ms at the 90th percentile, versus GPT Realtime-2.1's 860ms median and 1,900ms p90
- Word error rate of 6.9%, down from 7.8% without conversational context, ahead of dedicated ASR models in PolyAI's tests
- Turn-taking accuracy on caller interruptions approaches Gemini 3.1 Pro's performance
- English-only at launch; broader language support is planned
Dialog agents hear and handle a call the way a great human agent would - Matt Henderson, VP of Research at PolyAI
More AI news in Polish at nowosci.ai