Voice AI Is an Infrastructure Problem, Not Just a Model Problem

“Voice AI is an infrastructure problem because voice itself is an infrastructure problem,” says Kosta Pribić, a Senior Principal Engineer at Infobip.
He works with voice services – real-time communications across the globe, across multiple data centers, and for everyone reachable on the planet. Yep, it’s a lot.
That scale makes voice AI a much more complex challenge than developers might initially expect.
Developers can take control of the call in real time
Kosta sees voice AI as a two-layer system. Smaller models built into phones can handle real-time issues like packet loss and background noise. They should do their work quietly in the background, so the device just feels like it works.
Larger models, meanwhile, usually run in the cloud and need access to the audio being exchanged between devices. That’s where Infobip comes in: it provides the communications infrastructure developers need to process that audio in real time.
We give you control over the call. We enable you to interject, do something smart with the media.
He says a lot of the infrastructure behind this, from global data centers to enterprise deployments, was originally built for web pages. But web pages aren’t real-time communication.
So when teams use that infrastructure for voice AI, they end up carrying over decisions that were made decades ago. Those systems weren’t built for a live AI-assisted call, which means developers are often stuck working around constraints they probably wouldn’t choose if they were starting from scratch today.
Biggest challenges are real-time performance and cost
Developers entering voice AI often expect hallucinations to be the biggest challenge. But Kosta says they usually run into a different set of problems first, like custom signaling protocols, real-time delivery, and tricky networking edge cases. Networks aren’t built for this kind of traffic by default, so for developers without a voice background, it can feel like a completely new set of problems to learn.
“The main problem is understanding all of this. It’s completely new,” he says. That’s why Infobip’s first internally built LLM-enabled application was a voice AI troubleshooter. Kosta sees troubleshooting as a practical place to start, because teams first need to understand the communication problems they’re actually dealing with.
From there, model choice comes down to the business case and the margins. Kosta says a large model can make sense when latency isn’t a major concern and the economics work.
If your business case allows for using huge LLMs and you don’t care that much about latency, great, go for it.
But in most cases, he says, latency matters. Teams may need smaller LLMs to process calls, or fast, lightweight models to handle audio. Cost matters too, especially when a model is running on every call, because a large model can become expensive fast.
Voice AI has to be private, reliable, and real-time
People expect voice-enabled communications to be private and confidential, Kosta says. They also expect calls to be as reliable as the telecom services they already use, which he describes as operating at “five nines.” Once that trust is lost, it’s much harder to win back.
For developers who are new to voice, understanding those expectations is part of understanding the system they’re building.
That becomes especially clear in assistive use cases. Kosta points to an Infobip customer using a generative model to create voices for people with damaged vocal cords:
They can speak in a whisper while the person on the other end hears the generated voice. That’s, I think, the most heartwarming implementation I’ve ever done.
In the end, before choosing a model, developers need to understand how a call works, where it can break down, and how much delay and cost the application can handle. Even the best model won’t help if its response can’t make it back to the caller in time.


