Voice agents or voicebots come in various forms. The classic variant comes with significant latency, but technological development does not stand still. Speech-to-speech is faster and opens up new possibilities. Steam-connect immediately began exploring and offering this new voice solution.
Robin van Leyden is one of the two CTOs at Steam-connect – he focuses on R&D, voice solutions, and the UCaaS platform. “I have been working in the telecom sector since I was seventeen. In the early years, that involved separate telephone networks in office environments: pulling cables for internal networks.” More than 25 years later, the world of telephony has changed completely. The very latest: speech-to-speech voicebots.
Steam-connect is a supplier of both a UCaaS– as a CCaaS platform. The company has linked a voice agent to the UCaaS platform, which is active before calls reach the CCaaS platform. “With this, the voice agent plays a role in the 'zero line", as we call them," says Van Leyden.
Cascade and speech-to-speech – what is the difference?
Currently, there are two types of voicebots available. The first variant works according to the cascade model, with speech-to-text conversions, the use of an LLM, and text-to-speech conversion. The second variant works directly – speech-to-speech – and does not use text prediction, as with an LLM, but rather speech tokens. That technology is relatively new. Relatively, laughs Van Leyden, because: “The traditional cascade model hasn't been around for very long either.”
A well-known drawback of the cascade model is that the various conversions result in a relatively long processing time. Latency can rise to a second or more, and in certain situations, this is a real impediment to smooth interaction.
Van Leyden: “In that first model, ears, brain, and mouth – the various steps – are separate from each other. On top of that, all kinds of techniques are needed, such as VAD, voice activity detection, to determine whether speech is present and when the speaker has stopped. Those kinds of intermediate steps also require processing time. Companies invest a lot of time and money in the latency to reduce as much as possible.”
Voice tokens
"Speech-to-speech It has existed since 2024. In this direct variant, the ears, the brain, and the mouth are integrated into a single technology platform. The main advantage is that speech is and remains the source material, including everything contained within that speech. This includes intonation and emotion, speech characteristics that are lost as soon as you convert it to plain text. With speech-to-speech, that information remains available.
In speech-to-speech, speech is directly converted into tokens. The model reasons based on that speech information, uses the speech model to formulate an answer, and immediately generates a spoken response, Van Leyden explains.
Instead of a large text-based language model or LLM Speech-to-speech utilizes a large speech model. Large amounts of speech have been used as training material for this. This includes elements such as speed, volume, and emotion. The largest parties developing this technology are Google (with Gemini), OpenAI (with GPT-realtime), and xAI (with Grok). There is also an open-source model, Moshi, but it does not support Dutch.
Full duplex: talking and listening at the same time
The range of speech-to-speech solutions is currently still limited, which immediately makes it clear that we are still at the beginning of what is in itself a promising solution. This is partly because natural dialogues are taken to a higher level as the models are full duplex: listening and speaking, but even interrupting, is possible. Translation is also among the options. Most striking is the much lower latency, down to a few hundred milliseconds, according to Van Leyden.
(Text continues after the banner below)
For a demo of what full duplex entails: the Ziptone editors spoke with ChatGPT Live here.
At the same time, there are also some limitations. Observability—the recording of what happens during processing—is limited. Turn-taking is logged, and the tooling can also generate a summary. However, this lacks the text optimization provided by an LLM, says Van Leyden. “Therefore, it is recommended to run a speech-to-text (STT) solution and an LLM model alongside a speech-to-speech solution for transcriptions that form the basis for recording, summaries, quality monitoring, and analyses. Cascade models therefore certainly retain their right to exist.”
More human
Barge-in is another improvement in the dialogue with voicebots, according to Van Leyden.
“When a voicebot starts speaking, we as humans have a response ready very quickly. As soon as you start talking as a customer, the speech-to-speech bot stops speaking and starts listening, to then continue with the new information. This also makes it possible to generate interjections that confirm listening ('yes', 'OK', 'go on') or show understanding. Whether that is genuine understanding or more of a standard response, we haven't been able to figure that out yet. The technology is still brand new. But the voicebot that works with the cascade model cannot handle interruptions, let alone show understanding in the interim. This makes the new voicebot a lot more human.”
Learning to use a voicebot
There is something else we need to take into account: as humans, we must learn to interact with this second generation of voicebots. “We know from our own experience that consumers are used to speaking to automated speech systems in a dictation style or in a staccato manner—think keywords. With an input like 'problems with my phone,' a voicebot misses a lot of information that you could easily provide using full sentences, and thus with more depth. The bot could then, for example, search the knowledge base much better.”
Incidentally, there is a difference between open-ended question speech recognition, where the system primarily searches for recognizable keywords, and speech solutions that work with complete sentences based on LLMs, according to Van Leyden.
Should customer service managers educate customers on the use of new technology?
“People will indeed have to learn to communicate with these voicebots in a more natural way. On the other hand, under the AI Act, you have to point out that the customer is speaking to an AI system. For many people, a mental switch flips at that point: I’d better keep what I say simple. But AI is developing very rapidly, so consumers quickly fall a bit behind in terms of adoption and acceptance.”
Is speech-to-speech more data-intensive and therefore more expensive than the cascade model?
First and foremost: the voicebot performs – when necessary – a warm handover with call information to the agent, as implemented in the Steam-connect CCaaS platform, resulting in time savings. The billing model used by Steam-connect is a price per minute. This also includes LLM usage for transcription, which can be used for quality applications within the CCaaS platform. The cost per minute is always lower than the cost per minute of a human employee. The cost of a human agent – including employer contributions and all overhead – quickly approaches 2 euros per minute. Deployability naturally depends on the use cases and your knowledge base. We must not forget that a human agent regularly needs time to search for information on complex issues.
With the speech-to-speech voicebot, too, you will need to pay attention to setting up guardrails and escalation criteria. “These kinds of things are increasingly becoming a standard part of our customer contact work processes – something we will have to get used to. The technology will stand alongside us as a separate, smart application.”
Where do the greatest opportunities for speech-to-speech in customer contact lie?
Undoubtedly in support. The younger generation is communicating increasingly via voice messages on various platforms, so typing messages could well lose popularity in the coming years. For longer messages, I personally opt for voice as well. I expect voice solutions to help us more often, among other things by verifying whether they have understood us correctly – for example, before a voice message is sent. In the contact center, you can combine voice solutions perfectly with APIs that work with underlying systems. Think of checking the shipping status of a package. We may discover that certain matters, due to how we have currently set up processes, cannot be handled well with voice. Fortunately, you can always choose an escalation or a combination with text input.
“As Steam-connect, we have brought speech-to-speech in-house; it is now part of our platforms. For customers who want to get started with it, we offer a half-day session where we build a first application together. When it comes to standard interactions, that gets you a long way.”
“Speech-to-speech is still in its infancy, but it will play an increasingly large role in our lives in the near future,” is Van Leyden’s conviction. “The world is becoming a little smaller again, partly due to the possibility of direct translations. And the absence of annoying latency is a real advantage. Moshi, currently the leading speech-to-speech model in terms of latency, works end-to-end in 200 milliseconds. This brings us ever closer to having a normal conversation with an AI agent.”
(Ziptone/Erik Bouwer)
Corrections and additions: in an earlier version of this article, an indication of the rate per minute for the voicebot was provided. This has been removed in this version; Steam Connect does not yet wish to share rate information.
Featured, Technology


