Voicebots are being used increasingly often in customer contact. In this edition of Tech Update, we explain what voicebots are, what you can do with them, why the technology is relevant, and what the pitfalls are.
What is it?
A voicebot is a chatbot that communicates with the customer based on spoken language. Through natural language processing and machine learning, the voicebot can process what the customer says and respond via a synthetic voice. Another difference compared to chatbots is that a voicebot has two additional 'capabilities' on board: speech recognition to convert spoken language into text, and speech synthesis to provide generated answers in spoken form.
A voicebot initially operates without human intervention; the option to escalate a conversation to a human employee is ideally built in.
What can you do with it?
Voicebots can be deployed for various purposes. The task of the voicebot can be more or less limited and complex.
With open-ended question voice routing, the voicebot answers the customer's call, after which the bot asks the customer to state the reason for their call. Based on that input, the voicebot can transfer the customer to the correct representative. With this routing, provided the customer is identified or recognized, context such as subject, urgency, or customer characteristics can also be taken into account.
In addition to routing, a voicebot can also be assigned additional tasks based on the recognized customer intent. For example, assisting the customer further by providing them with a link via SMS to a page containing a suitable self-service solution. Another possibility is that the voicebot answers questions directly based on a further dialogue with the customer. When the voicebot operates in a private environment (for example, the customer's 'my account' area), it can offer personalized service by combining the knowledge base with customer data.
Use cases of voicebots
- Identifying the customer request and ensuring the right solution: referral to self-service, automated processing, or forwarding to an employee;
- executing payment orders within a mobile app, answering questions about expected debits or spending patterns, bank cards block or unblock;
- assistance with mortgage and credit applications, such as collecting documents and reminding of deadlines);
- providing real-time information about policy updates and payment statuses of insurance claims;
- appointment logistics, scheduling follow-up appointments, data collection and lead qualification;
- settle service requests for car dealers, returning missed calls;
- Drive-thru assistant for fast food restaurants for taking orders or making upsell suggestions.
Why is the voicebot relevant?
By far the majority of contact centers are primarily set up for communication via voice. While this is the most effective channel, it is also the most expensive. When voice-based communication between customer and company can be automated, it offers significant savings potential. A good (voice) bot can take a lot of work off your hands, but is also available 24/7, is scalable, can respond quickly, and offers inclusive service, partly because a voice bot can assist in various languages.
Although the building blocks of the voicebot – speech recognition, speech synthesis, and chatbot technology – have existed for quite some time, the application of voicebots is still in its infancy. However, it is expected that voice will become increasingly important for consumers to interact with automated systems as well.
How does it work?
A customer's speech is converted into text via speech recognition, after which the intent is derived from the text, with or without the use of other information. The application then formulates a text response based on knowledge bases and other information. Synthetic speech software converts the text into speech, whereby the voice type and intonation can be customized.
Advanced voicebots are capable of translating what the customer wants to achieve (the intent) into commands for AI agents. In doing so, the voicebot remembers the goal and the context and can intervene in the event of errors or freezes.
Some software providers package the various conversion steps into a single package, such as Microsoft's Voice Live API, where speech recognition, generative AI, and text-to-speech are integrated into a single interface. The latest developments point towards native audio or speech-to-speech models, such as OpenAI Realtime and Google Gemini Live/Native Audio, which process audio directly and immediately provide an audio response. This method has the advantage of a lower latency is, but with the disadvantage that no transcript is required. A transcript is often needed not only for multiple conversion, but also for logging, searching, analysis, editing, linking to knowledge bases, and use for intent recognition, quality assessment, and reporting.
The advantages of direct conversion are that a more natural conversation can emerge and that factors such as intonation and tempo continue to play a full role. At the same time, there are strong indications to date that end-to-end speech-to-speech models deliver weaker output than solutions that make the intermediate step of text. For the time being, companies will likely continue to opt for the step-by-step conversion from speech to text and back again.
Just pay attention
1. The application of voicebots is still in its infancy. This is due to both technical and human aspects of communication.
2. Voice-based conversation presents additional challenges compared to a text-based chatbot. For instance, dealing with human speech: people may not finish their sentences, change the subject unexpectedly, or may not always be clearly understandable due to various reasons. Listening to speech and converting it into text takes time (the aforementioned latency), meaning a voicebot's response does not always follow immediately after a customer has said it. Factors such as ambient noise or connection quality can also cause performance issues.
Latency – De latency (The delay between capturing the customer's speech and subsequently responding with speech from the voicebot) is caused by various types of steps and conversions, including waiting until the customer has finished speaking, processing the speech, searching knowledge bases, formulating a response in text, and converting that text into speech. Currently, the average latency is between 0,7 and 1,0 second. Whether this latency is bothersome depends heavily on the context and the application. A latency of around half a second is acceptable to many people, and a latency of 200 to 300 milliseconds is no longer noticed.
The challenge, therefore, lies not in speech synthesis, but in gaining a good understanding of what the customer is saying. Additionally, voicebots primarily draw from the same information sources as chatbots, so the same applies here: the bot must have access to a high-quality knowledge base. This is already relevant for simple use cases, such as routing based on recognizing spoken intent. The speech recognition software must then, for example, be familiar with the specific jargon of the customer and the company.
3. Some consumers find voicebots less appealing than chatbots. Customers may be bothered by poor audio, latency, or accents. Additionally, the chatbot's understanding of the content may fall short. Users, for example, may end up in a loop. Aside from comprehension issues, many consumers are also concerned about what is stored and/or shared by a voicebot. Experts also point to the risk of reputational damage from inauthentic AI voices and to problems that can arise if backlogs in knowledge base maintenance occur.
What do you think: voicebots in customer contact, hot or not?
Hot or not: voicebots in customer contact
- HOT (100% 1 Votes)
- NOT (0% 0 Votes)
Total Voters: 1
(Ziptone/editors)
Featured, Knowledge partners, Technology, tech update



