Last Updated on August 5, 2026 by Craig Allen Keefner
AI voice order recognition is revolutionizing user interaction in self-service kiosks by introducing voice-enabled interfaces as an alternative to traditional touchscreens. This technology aims to enhance accessibility and convenience, driven by the increasing comfort users have with voice commands from home automation devices. However, achieving high accuracy in real-world, noisy environments, such as a McDonald's where noise levels can reach 80 decibels, remains a significant challenge. Kiosk Industry notes that current prototypes often rely on human assistance to ensure 99% order accuracy, highlighting accents, dialects, and ambient noise as key hurdles for fully automated AI voice systems.
McDonald’s 2024 Prototype (with listeners)
Kiosk Voice response promises to add new interactivity for self-service devices
how do kiosks help people with speech needs?
Kiosks can help people with speech needs by offering voice-enabled interaction, providing an alternative to traditional touch-based interfaces. While the text block primarily discusses the broader shift to voice interaction in self-service devices, driven by technologies like Amazon's Echo and Google's Home, the underlying principle of voice input can be adapted to assist users with various communication challenges. A growing number of technology vendors are introducing voice-enabled kiosks, exploring how interactive voice response can fill new needs in interactive kiosks, including potentially aiding those with speech difficulties by offering a different mode of communication.
Over the past few years, though, the concept of interactivity has taken on a new dimension. Driven in part by home automation devices such as Amazon’s Echo and Google’s Home, people are becoming increasingly comfortable with a new way of interacting with self-service devices: by voice.
A growing number of technology vendors have been introducing voice-enabled kiosks over the past few years. The question remains, though: what does the future hold for interactive voice response and what needs will it fill when it comes to interactive kiosks?
McDonald’s Voice Order
Nice video of prototype McDonalds Voice Order. Imagine 3 of these side by side in a NYC McDonalds. Ambient noise level in a restaurant can easily hit 80 db (Noisy Planet NIH). This prototype Kiosk ordering has 99% accuracy because a human is in the loop on every order…
Accents, dialects and languages are TBD but here is response.
It’s a nice demo albeit not a real restaurant with the typical ambient noise.
The usual questions regarding accents, dialects, languages along with noise come into play.
More Info
The average noise level in a McDonald’s restaurant can vary, but it is generally quite loud. According to a report by Noisy Planet, noise levels in restaurants can average 80 decibels (dBA) or higher, which is significantly louder than a typical conversation at about 60 dBA1. This level of noise can make it difficult for patrons to have conversations and may even pose a risk to hearing over prolonged exposure. It’s recommended to use earplugs or earmuffs in loud environments to protect your hearing.
What challenges are slowing the adoption of voice recognition kiosks?
Simply put, an interactive voice response system is a computer interface that accepts input by voice rather than mouse, keyboard or touch. The technology has been around at least since the 1970s but has become increasingly widespread as large organizations deploy such systems to handle customer service. And when combined with artificial intelligence, it’s becoming increasingly difficult to distinguish VR from communication with a live person.
When it comes to self-service kiosks, a quick Internet search shows dozens of vendors offering devices outfitted with a VR interface. Such interfaces are touted as a way to provide access for those with limited hand mobility as well as those who can’t read. As is the case with on-screen touch menus. It’s relatively easy to incorporate a variety of languages into VR, allowing the deployer to serve those with a limited command of English.
The adoption of voice recognition kiosks is primarily slowed by the need for new hardware, specifically the inclusion of microphones and speakers in existing kiosks, and environmental noise considerations. Rob Carpenter, CEO of Valyant AI, highlights that the biggest hindrance is the lack of embedded microphones and speakers in past hardware iterations. Additionally, the environment where the kiosk is located is crucial, as conversational AI can struggle in high-traffic, noisy areas like airports. Deployers must consider hardware capabilities to handle conversational AI and design considerations like microphone arrays and noise-absorbing materials. While interactive voice response technology has been around since the 1970s and offers benefits like accessibility for those with limited hand mobility or who can't read, these hardware and environmental factors are key obstacles to widespread implementation.
“Voice recognition is ready for kiosks and companies like Zivelo are already looking at ways to begin rolling the technology out on a wider scale,” said Rob Carpenter, CEO and Founder of Valyant AI, an enterprise-grade conversational AI platform for the quick-serve restaurant industry.
“The biggest hindrance to adoption and scale is going to be the inclusion of microphones and speakers in kiosks, which are required for conversational AI, but hadn’t been included in past hardware iterations because they weren’t needed at the time,” Carpenter said.
The environment where the kiosk will be located will also be a consideration.
“It’ll be important to look at the hardware’s ability to handle conversational AI (it’ll need embedded microphones and speakers), but it’s also important to consider the noise level in the environments,” Carpenter said.
“Conversational AI might struggle in high traffic areas like airports where there is so much noise it’s hard for the AI to hear the customer,” he said. “It’s very likely that for the highest and best use of conversational AI in kiosks, it may also require other capabilities like lip reading and triangulating the customer in a physical space to separate out disparate noise channels.”
As such, deployers will need to incorporate design considerations that include microphone arrays focused on specific areas where a user might be standing. They’ll also need to incorporate design considerations beyond the kiosk itself, including noise-absorbing carpet and walls in the area where the device will be located.
What are the privacy concerns with voice-enabled kiosks?
Privacy concerns will come into play as well. Amazon’s Echo devices, for example, store a record of what they hear when activated. And while such recording is only supposed to occur when the user says a “wake” word such as Alexa, anyone who owns such a device knows similar words can prompt a wakeup as well. In addition, when someone is using a VR-enabled kiosk there’s a distinct possibility that nearby sounds will be picked up and recorded as well.
“[It’s a concern] not only for the person ordering train tickets, but for the person who might be standing next to that person who’s having a quite high-level conversation on the phone with a business colleague—or his mistress,” said Nicky Shaw, North American distribution manager with Storm Interface. Storm designs, develops, manufactures and markets heavy-duty keypads, keyboards, and custom computer interface devices, including those that provide accessibility for those with disabilities.
“Now that’s also been picked up and sent to the cloud,” she said. “Privacy needs to be given more consideration in my view because just deploying a microphone on a kiosk with no visible or audible means of letting people know it’s always on needs to be factored into the design.”
What are the accessibility protocols for voice-enabled kiosks?
The protocols and practices for implementing voice in kiosks are not addressed in any U.S. Access Board standards and the KMA with Storm have incorporated a proposed voice framework for accessibility and more. The Access Board has these standards to consider as a baseline for when they create actual standards. In that sense KMA is setting the table for them.
The degree to which companies mine voice data for advertising information creates its own set of privacy concerns. Because most voice user interfaces require cloud processing services, any time the voice leaves the device makes the process more susceptible to a privacy breach.
Currently, the protocols and practices for implementing voice in kiosks are not addressed in any U.S. Access Board standards. However, the Kiosk Manufacturer Association (KMA), in collaboration with Storm, has incorporated a proposed voice framework for accessibility, aiming to set a baseline for future official standards. Beyond formal protocols, the success of voice-enabled kiosks also hinges on ease of use for the average person. While voice input is the collection method, the processing machine learning technology needs significant improvement, as noted by Tomer Mann of 22Miles, who states, "We are moving forward with integration but there is a long way to go." The text also touches on related privacy concerns regarding voice data mining and potential branding confusion, which are indirect considerations for the overall implementation and acceptance of voice accessibility.
And at the end of the day, making it easy for the average person to use will go a long way toward determining how successful VR in interactive kiosks will be.
“Voice input is the collection method, while the platform collecting the command is the brain/processing power to take the correct actions,” said Tomer Mann, EVP for Milpitas, Calif.-based software company 22Miles.
“We are moving forward with integration but there is a long way to go,” Mann said. “We have the input command solution but the processing machine learning technology needs to improve. It will happen with a few more iterations and innovation.”
What are the applications and impact of voice recognition in self-service kiosks?
One of the obvious applications for VR in self-service kiosks is for accessibility, enabling their use by those with impaired vision or limited hand mobility.
VR can also be used to create the “wow” experience business operators are looking for. Imagine, for example, the opening of the latest blockbuster superhero movie.
“Let’s say a video wall at the theater senses that someone is approaching,” said Sanjeev Varshney, director, Global SAP with Secaucus, N.J. based Cyntralabs, a developer of integrated solutions that help retailers drive sales.
“It could display a character from the movie, who says something such as ‘what movie would you like to see?’,” he said. “The character could then point to a card reader and say ‘just insert your credit card here” and have the tickets printed out or have an SMS sent to your phone.”
“One driver for voice relates to efficient and faster transactions” said Joe Gianelli, CEO & cofounder of Santa Cruz, Calif.-based Aaware Inc., a developer of technology that enables voice interfaces.
Consider tasks that may require an excessive amount of screen navigation or drilling down, Gianelli said. Voice is usually much more efficient if the user needs to navigate beyond three levels of touch.
Of course, VR won’t be a catch-all solution. Still, VR could be part of a menu of accessibility options.
“Speech command technology will never replace the need for other interface devices because people with speech impediments won’t be able to use it, just like there are people who are blind and can’t use a touchscreen,” Shaw said.
“A deployer would still need to provide tactile interface devices as well as the speech command,” she said. “This needs to be seen as another element in multimodal accessibility. There’s not a one-size-fits all solution.”
The technology is at its infancy, but with further innovations and feature updates, the solutions will only be more agile to day-to-day user experiences,” Mann said.
“Technology is getting there,” he said. “22Miles just wants to stay ahead of that innovation as we do it all other digital or content triggering capabilities.”
And when it comes to industries, some of the key applications insiders are seeing are in the ticketing and restaurant ordering fields, with initial results showing promise. Catalogue lookup in a retail setting might also be a prime candidate.
“Imagine being able to find, filter and sort any item through voice,” Carpenter said. “It would eliminate the tedious tasks of searching through pages and pages of items to find your favorites. Just tell it what you want and then be on your way.”
More Information
WHITEPAPER – VOICE RECOGNITION & SPEECH COMMAND ASSISTIVE INTERFACE
MASTERCARD ZIVELO VOICE ORDERING WITH AI
KROGER LAUNCHES VOICE ASSISTANT ORDERING FOR GROCERY ECOMMERCE
ALEXA SELF-ORDER VOICE COMMAND VOICE RESPONSE QSR W/ CUSTOMER & EMPLOYEE. BEACON TECH
Frequently Asked Questions
- 1 What are the primary challenges associated with deploying AI voice order kiosks in quick-service restaurants?
-
The page highlights significant challenges such as high ambient noise levels in restaurants, which can impact voice recognition accuracy. The McDonald's 2024 prototype, for instance, currently relies on human assistance for 99% accuracy, indicating limitations in fully automated systems. Addressing diverse accents, dialects, and languages also remains a key consideration.
- 2 How does AI voice order technology enhance user interaction in self-service kiosks?
-
AI voice order introduces a new dimension of interactivity beyond traditional touch-enabled displays, leveraging the growing user comfort with voice commands from home automation devices. It promises to add new ways for users to interact with self-service devices, potentially making the experience more intuitive and accessible for some.
- 3 What privacy and accessibility considerations are important for AI voice order kiosks?
-
The page's table of contents explicitly identifies "Privacy Concerns" and "Accessibility Protocol" as critical factors impacting the adoption of voice recognition kiosks. These aspects are crucial for ensuring user trust and compliance, and are noted as considerations slowing the technology's widespread implementation.
- 4 What are some current examples or prototypes of AI voice order technology in kiosks?
-
The page references the McDonald's 2024 Prototype, which includes voice listeners for ordering. Other mentioned examples include Mastercard Zivelo Voice Ordering with AI, Kroger's launch of voice assistant ordering for grocery ecommerce, and Alexa self-order voice command systems. These illustrate ongoing developments and applications in the industry.