Testing Android’s On-Device Speech Recognition With Hindi-English Code-Switching

Wait 5 sec.

My mother describes how she feels in Hindi, with English health words dropped into the middle of the sentence. "BP high hai." "Fever lag raha hai." "Dizzy ho raha hai." Nobody in my family thinks of this as code-switching. It's just how we talk.I'm building an Android app so she can log how she feels by voice instead of typing. Her village connection is unreliable, so the app has to work offline, and because this is health data, I didn't want her audio going to a server if I could avoid it. That means on-device speech recognition, and that's where I found out how many things can quietly go wrong between "the API exists" and "my mother is understood".This is what I learned, including the parts that didn't work.The setupAndroid has SpeechRecognizer.createOnDeviceSpeechRecognizer() (API 31+), which runs recognition locally with no network. I use it live: she taps the yellow "something hurts" button, speaks for up to 15 seconds, and the transcript is saved with her check-in. No audio is kept. The text is what a doctor will later read.Step one is asking for Hindi:putExtra(RecognizerIntent.EXTRA_LANGUAGE, "hi-IN")putExtra(RecognizerIntent.EXTRA_PREFER_OFFLINE, true)That worked well enough that I almost stopped there.The test that mattered: a real speakerI could have tested by playing recordings into the phone's mic. I didn't, because loudspeaker playback plus room noise tends to understate real accuracy, and I'd be testing my own assumptions. I sat down with my mother and had her say the sentences she'd actually say, mixing in English the way she does, and I wrote down what she said next to what the phone heard.The result was mixed. Some English words came through fine. "Treatment area" and "walk" were transliterated correctly into Devanagari. But "dizzy" came back as डीसी, which is the letter names D-C, not the word.For a health log that's a nasty failure. It doesn't look like an error. It looks like a plausible Hindi word. A doctor reading डीसी has no way to know it was supposed to mean dizzy.The fix was one flag and one trade-offAndroid 14 added language switching to the recognizer: you list the languages it may move between, and it follows the speaker. I'd enabled it with LANGUAGE_SWITCH_BALANCED, which favours speed.if (Build.VERSION.SDK_INT >= Build.VERSION_CODES.UPSIDE_DOWN_CAKE) { putExtra( RecognizerIntent.EXTRA_ENABLE_LANGUAGE_SWITCH, RecognizerIntent.LANGUAGE_SWITCH_HIGH_PRECISION, ) putExtra( RecognizerIntent.EXTRA_LANGUAGE_SWITCH_ALLOWED_LANGUAGES, arrayListOf("hi-IN", "en-IN"), )}Switching to LANGUAGE_SWITCH_HIGH_PRECISION fixed "dizzy". In a retest with the same speaker, "BP", "fever" and "tablet" came back right too.The cost is latency. The recognizer takes longer to commit to a language. I accepted that on purpose: this is a 15-second check-in that a doctor reads later, so a slightly slower answer that's correct beats a fast answer that turns a symptom into two letters.I wrote the reason in a code comment with the date, because in six months I'd otherwise "optimize" it back.A limit: I can only claim what a real speaker checkedI validated Hindi with her. English (en-IN) needs its own real-speaker test, and it isn't claimed as validated in my submission. The rule I set: no language goes in the submission unless someone who speaks it has tested it. It's slightly embarrassing to ship with fewer languages than I'd like. It's much better than shipping a claim I can't stand behind.I also nearly made one mistake here. I considered pairing an English profile with Hindi, so Hindi words would come out in Devanagari. I dropped it the same day, before any build shipped it. A person who speaks Indian-accented English and no Hindi could have their English recognized as Hindi, giving unreadable output, which is worse than Hinglish written in Latin letters. The idea was never tested with that kind of speaker, so it stayed out.A phone that throws on purposeA OnePlus 9R running Android 12 crashed the app. Android provides isOnDeviceRecognitionAvailable() from API 31, but I wasn’t checking it before calling createOnDeviceSpeechRecognizer(). The latter throws synchronously when the phone has no on-device recognizer, and I hadn’t wrapped it either:val speechRecognizer = try { SpeechRecognizer.createOnDeviceSpeechRecognizer(context)} catch (e: UnsupportedOperationException) { knownUnavailable = true trySend(RecognitionState.Failed(RecognitionFailure.UNSUPPORTED)) close() return@callbackFlow}Now it fails once, remembers it, and every later check-in on that phone goes straight to recording-and-upload instead of crashing. If you're supporting Android 12 at all, catch this.Let the user veto the transcriptSince on-device mode keeps no audio, a wrong transcript would be unrecoverable. So there's a "That's not what I said" button. It deletes the transcript and re-records the entry through the cloud route. The rejected text is not kept next to the correction. The entry exists to record what she actually reported, and a sentence she's disowned shouldn't reach a doctor.What I'd tell another builderTest with a real speaker and write down word-for-word what was said versus what was heard. The gap between a correct transliteration and a plausible wrong word is the whole result.Check what a device supports before you design around it. On Android 12, the on-device recognizer may not exist at all.Code-switching isn't an edge case in many households. If your users mix languages, test that first.Wrong-but-plausible output is more dangerous than an obvious error. Give the user a way to reject it.CareVocal is my Shipaton 2026 entry, built for my mother and logging only, never diagnosing.