Engineering

Uzbek speech recognition in production: what works out of the box and what we had to push

Voice AI in Uzbek is bottlenecked not by the model but by data and latency. We break down where the milliseconds and accuracy leak — and how we got them back.

Engineering teamZukko.AI ML & infrastructureJune 17, 20269 min read
ASR · LATENCY

A voice agent for a contact centre is not “a model that dials people”. It is a pipeline: speech recognition, request understanding, response and synthesis — all of which must fit within a latency where a human feels no pause. In Uzbek the bottleneck is recognition itself: accents, switching to Russian mid-sentence, banking terminology.

What works out of the box

The base model handles clean studio speech and standard phrases well. For FAQ scenarios — “card balance”, “opening hours”, “branch address” — that is nearly enough. The trouble starts on the phone channel: compressed audio, noise, interruptions.

What we had to push

We raised accuracy on domain vocabulary not by swapping the model but with data: labelled real calls, a glossary of terms, handling of “Uzbek ↔ Russian” code-switching. Latency was a separate fight — streaming recognition instead of waiting for the end of a phrase removes the most noticeable pause.

In voice AI it is not the largest model that wins, but the shortest, most predictable path from sound to answer.

Takeaway for anyone choosing a voice vendor

Ask not “which model do you use” but “on what data is it fine-tuned for our language and our industry” and “what is the latency on the phone channel”. For CIS banks those are the two decisive questions. Exact figures for your environment are confirmed on a demo.

DEMO

Launch an AI employee on your own channels

We will assemble a pilot on your real chats from Instagram, Telegram and telephony — you will see the result on your own numbers. SMB launch from 2 hours.

Uzbek, Russian and English. Connects to your channels and systems.