Why AI gets weaker when your child asks in Hindi or Tamil

A chat bubble motif for a question asked in a home language

A child sitting down to homework types the question the way she thinks it. Half Hindi, half English, in Roman letters, because that is how the phone keyboard works and how her friends message. The answer comes back instantly, fluent and confident.

That confidence is the problem. When a model gets weaker, it does not sound weaker.

The gap has actually been measured

Researchers at the Nilekani Centre at AI4Bharat, IIT Madras and IBM Research India built a benchmark called MILU to test exactly this. They scraped tens of thousands of multiple-choice questions from real national and state exam papers in India, covering 11 languages, eight domains and 41 subjects, then ran more than 40 large language models through them.

The best performer was GPT-4o, at around 74 percent on average. Split by language, though, the same model scored about 82 percent on the English questions and about 70 percent in Tamil. The paper’s own summary is blunt: while many models claim to support multiple languages, “there is still a huge discrepancy in their performance in English and other languages.”

The pattern inside those numbers matters more than the headline. Models did worst in the culturally specific subjects, arts and humanities, law and governance, and best in general STEM, which the authors put down to training data that “lack sufficient culturally specific data”. The weakness is not spread evenly. It concentrates exactly where a child’s history, civics and literature homework lives.

One honest caveat: MILU was published in late 2024 and revised in early 2025, and models have improved since. Nobody has published evidence that the gap has closed.

Script and voice add their own losses

Then there is how the question gets typed. A December 2025 study looked at 3,153 real messages sent to a maternal health helpline in India, in Hindi, Telugu, Kannada, Marathi, Punjabi, Nepali and English, and compared native script against romanised text. Most models did worse on the Roman version. GPT-4o lost only about 2.6 points, but some open models lost seven or eight, and for Nepali the drop reached the twenties. The researchers also found romanised inputs “exhibit higher flip rates”, meaning a trivial spelling change flipped the model’s answer while the meaning stayed identical.

That study is about medical triage, not homework, so treat it as a signal rather than a settled finding. But the direction is worth knowing, because romanised typing is the normal case for children.

A hand typing on a phone keyboard
Most children type their home language in Roman letters, which is the input models handle least consistently.

Voice input has the same unevenness, and it is easier to check. In February 2026, AI4Bharat at IIT Madras and Josh Talks published Voice of India, a speech-recognition benchmark built on 536 hours of audio from 36,691 speakers across 15 languages. On its leaderboard the spread between systems is enormous: Google’s Gemini 3 Pro posts a word error rate near 6 percent, while OpenAI’s GPT-4o Transcribe sits around 34 percent. Same child, same sentence, wildly different transcript depending on the app.

Meanwhile, most children think it is a search engine

The Central Square Foundation surveyed 15,000 parents, children and teachers across 10 Indian states between August 2025 and January 2026 for its Bharat Survey for EdTech, published in February 2026. Among children already using edtech, 35 percent use generative AI for learning, and 69 percent of those use it daily.

Here is the line that should give a parent pause. Of the children who said they understood generative AI, three-quarters confused it with internet search.

Put the two findings side by side. A child who believes she is looking something up, asking in a language where the model is measurably weaker, in a script it handles least consistently, has no reason to check anything. Search engines are not supposed to guess.

None of this is an argument for switching the household to English. India is going the other way on purpose: ahead of the India AI Impact Summit in New Delhi this February, the government-backed BharatGen consortium said its Param2 model would cover 22 Indian languages, with its chief executive Rishi Bal saying that “India’s AI progress must be built on language access and contextual understanding.” That is the right direction. The habits just need to arrive first.

Four things to do at the homework table

Split the job. Let the child ask for the explanation in the language she thinks in, because understanding is what the home language is for. Then check the facts separately: names, dates, formulas, definitions, either in English or against the actual textbook.

Type in the script when the answer matters. Roman letters for chatting is fine. For a question whose answer will be written into an exam paper, the proper script is the safer input.

Catch it on home ground, once, together. Ask it something about your own district, your own festival, your state board’s syllabus. That is the ground where the benchmarks say models are weakest, so a wobble is likely, and a child who watches the machine fumble a fact she already knows learns more in that minute than in an hour of warnings.

Read the answer out loud. If the Tamil or the Hindi sounds like nobody’s grandmother, stiff, translated, textbook-ish, that is a clue the model produced an English answer and converted it. Converted answers carry over English framing, and that is where the cultural details go missing.

The read from Singapore

This is not only India’s problem. SEA-HELM, the evaluation suite built by AI Singapore with NUS and published in 2025, covers Filipino, Indonesian, Tamil, Thai and Vietnamese, and includes Tamil specifically because it is an official language of Singapore. Its finding is the same shape: significant gaps remain between English and these languages, with Tamil among the weakest performers.

So a Tamil-speaking family here faces the same accuracy gap as a family in Chennai, without a national programme aimed at closing it. We build AI mentoring software, so we have an obvious commercial interest in parents trusting AI with their children’s learning. That is precisely why this is worth saying plainly: the fluency of an answer in your language tells you nothing about whether it is right.

The gap will narrow. Better data and benchmarks like these exist to close it, and they probably will. What will not go out of date is a child who reaches the end of a confident paragraph in her own language and thinks to ask how she could check it.

A notebook and pencil on a study desk
The checking habit costs about a minute and outlasts every model upgrade.

Related reading: how to teach your child to fact-check an AI, and why Singapore’s school AI listens to kids instead of answering.

Was this useful?

Sources

  1. MILU: A Multi-task Indic Language Understanding Benchmark · arXiv (AI4Bharat, IIT Madras and IBM Research India)
  2. Script Gap: Evaluating LLM Triage on Indian Languages in Native vs Romanized Scripts in a Real World Setting · arXiv
  3. Voice of India speech recognition benchmark · AI4Bharat (IIT Madras) and Josh Talks
  4. Beyond Access: Insights from the Bharat Survey for EdTech 2025 · Central Square Foundation
  5. SEA-HELM: Southeast Asian Holistic Evaluation of Language Models · arXiv (AI Singapore and NUS)
  6. BharatGen to launch 17 billion parameter multilingual AI model at India AI Impact Summit · Business Today