Consider a fairly common scenario. Someone who got their first smartphone years ago and has barely typed a few hundred words on it since. Everything runs through voice. Long WhatsApp voice notes. YouTube search by speaking. And the phone gets spoken to the same way any person would be spoken to, in Hindi, with English words dropped in wherever Hindi starts feeling like too much effort.
"Sharma aunty wala paneer recipe nikaal do."
That sentence used to be a genuine problem for voice systems. The phone would catch the Hindi and mangle the English, or catch "recipe" and give up on the rest. The command would then get repeated more slowly, which rarely helps, and eventually the phone would be handed to whoever was sitting nearby. This played out for years, across countless households.
It largely works now. Which sounds like a small development, and isn't.
Voice isn't a convenience feature here
A lot of product thinking carries an underlying assumption that voice is what gets used when hands are busy. Driving, cooking, that sort of situation. A nice-to-have layered on top of the real interface, which is the screen.
That framing doesn't hold up in the Indian context.
The people coming online now are not people who will eventually learn to type and then stop relying on voice. Someone who speaks a regional dialect at home, reads Devanagari slowly, and has never opened a Hindi keyboard layout isn't going to switch to typing. Typing in one's own language on a phone remains genuinely difficult, which is why a large share of the country types Hindi, Tamil and Bangla using English letters and hopes it's understood on the other end. Voice bypasses that entire problem.
Add earbuds into the picture, where there's no screen at all. On a device with no display, voice isn't one option among several. It's the only way in. Whether the voice model performs well or poorly more or less decides whether the product feels pleasant or frustrating. There's no fallback to lean on.
Why this is harder than simply adding more languages
This part tends to be underestimated, including by teams building these products.
Take Hindi as an example. On paper, it's one language. In practice, Hindi spoken in Lucknow, Patna, Indore and Jalandhar involves fairly different acoustic patterns, and a model trained mostly on clean, Delhi-accented Hindi will struggle elsewhere. Extend that across every major language in the country, and the real scale of the problem becomes clear. It isn't twenty-two languages. It's a few hundred ways of speaking twenty-two languages.
Then there's language mixing, which happens constantly and without warning. "Bhai volume thoda kam kar, meeting hai." Two languages in one sentence, with the switch happening mid-clause. Older systems expected users to select a language in settings and stay within it. That isn't how people naturally speak. Solving for this requires models trained on how language is actually used, not two clean monolingual datasets stitched together.
Data availability is another significant constraint. English has decades of transcribed, labelled audio available. Many Indian languages have very little usable data, and some have almost none. This is the unglamorous reason initiatives like Bhashini and the work coming out of AI4Bharat matter more than they get credit for. Someone has to build that foundational layer first. Without it, every company would need to collect data from scratch, which isn't commercially viable for a language spoken by a few million people. The economics simply don't support it.
Names are a smaller but frequently overlooked issue. Station names, dish names, personal names. These trip up recognition systems more often than ordinary sentences do. "Chalo Thiruvananthapuram" or "Debjani ko call lagao" can break recognition faster than most standard commands.
What actually changed
Largely, the architecture.
The earlier approach worked in stages. Audio to text, text to intent, intent to response. Each stage introduced its own errors, and those errors compounded, resulting in systems that were technically functional but practically frustrating to use. Newer end-to-end models learn directly from audio, and turn out to be significantly more tolerant of accents, incomplete sentences, and mixed-language speech.
The other major shift is that basic commands now run directly on the earbud itself. "Next song," "volume badhao," "call kaato" happen locally and instantly, with no server involved. Heavier queries still route through the cloud. Anyone who's had earbuds go unresponsive inside a Metro tunnel has experienced that exact gap.
Where Indian brands come into this
For a long time, the assumption was that the serious voice work would come from global players, and Indian brands would simply license whatever trickled down. That's starting to change, and boAt's Crest AI is a useful example of the shift.
Crest AI is boAt's own proprietary AI platform, built to sit inside its next generation of earbuds and headphones. The idea is to move the earbud beyond music and calls into something closer to an always-available companion. One that can be spoken to naturally, asked questions, used to pull up real-time information, get things translated or explained, and offer assistance that's actually personalised to the user. All through voice, without reaching for the phone every two minutes.
What makes it relevant to this particular discussion is that it's being built for Indian speech from the ground up. Crest AI is designed to understand and respond across multiple Indian languages and dialects, which is precisely the problem this entire article has been describing. A household that switches between Hindi and English mid-sentence, a user whose accent doesn't match the training data, a command that includes "Thiruvananthapuram" without warning. Global assistants have historically treated all of this as edge cases. For a platform built here, it's the main case.
There's a sensible logic to a homegrown audio brand doing this. boAt already ships earbuds at a scale most global brands don't manage in India, which means the hardware, the software and the AI layer can be designed together rather than bolted onto each other. Whether the execution fully lives up to the ambition is something real-world usage will decide, and the same tests suggested below apply to it as much as to anyone else. But the direction is the right one. An AI companion in the ear, speaking the languages people actually speak, is a far more natural fit for this market than another assistant that works beautifully in polished English and falls apart at "meeting hai."
What to check before buying
Packaging often states something like "supports 10 Indian languages." That claim is worth testing directly in the language actually spoken, rather than taken at face value.
A useful test: deliberately give a command that switches languages midway through. A system that handles this reasonably well is a strong one. A system that requires committing to a single language before speaking is an older system with newer marketing.
It's also worth avoiding showroom-only testing, since showrooms tend to be quiet environments. Testing outdoors, near traffic, gives a far more realistic picture. Voice recognition tends to break down in noise, which is why the mic array and noise separation on the device matter just as much as the underlying language model. These are often treated as separate specs when they really aren't.
Checking what functions without internet is equally important. Volume, playback and calls shouldn't require a network connection. If they do, the system is largely cloud-dependent, and that becomes noticeable the moment signal drops.
One more detail worth flagging. Some devices understand a spoken language correctly and then respond in English regardless. Technically functional. Practically, not quite the intended experience.
It's still patchy
Coverage remains uneven. Widely spoken languages are served reasonably well; smaller languages and most dialects are not, and that gap isn't likely to close quickly. Accuracy drops in noisy environments and drops further for older speakers, whose speech patterns remain underrepresented in most training datasets. Fast speech continues to cause issues, and cloud response delays can be long enough that users simply give up and reach for the phone instead.
Privacy is a legitimate consideration here too, not an exaggerated one. A device listening for a wake word is, by definition, listening. And a device built to be conversed with, whether it's a global assistant or a platform like Crest AI, carries the same question. What gets stored, what gets transmitted, and how that can be switched off should be easy to find within the app, ideally within half a minute. If it isn't, that's worth noting.
At the end of the day, most users never think about whether their sentence stayed grammatically consistent across two languages, and they shouldn't have to. People simply speak the way they always have, and the system either keeps up or it doesn't.
For most of the last decade, it largely didn't. That's the part that's changing now, and for once, some of that change is being built at home.
