The focus of Indian language AI research is on rigorous, real-world evaluation. Recent preprints address challenges in automatic speech recognition (ASR) for diverse Indian languages and the complexities of code-mixed Hinglish, particularly in detecting harmful content and hallucinations. These efforts introduce new benchmarks and diagnostic tools to measure AI performance in the messy, varied conditions of actual use.
Vimarsha: Realistic ASR Evaluation Across 22 Indian Languages
A new benchmark called Vimarsha, submitted on September 21, 2026, by authors Kaushal Santosh Bhogale, Srija Anand, Sadakopa Ramakrishnan Thothathiri, Tahir Javed, Sshubam Verma, and Mitesh M. Khapra, corrects systemic biases in Indian language ASR evaluation. Existing benchmarks, according to the researchers, tend to produce optimistic scores due to their reliance on clean, controlled audio. They also yield pessimistic scores by penalizing valid linguistic variations with overly rigid transcription standards.
Vimarsha is a 100-hour benchmark that covers all 22 scheduled Indian languages. It incorporates demographically diverse on-field recordings and in-the-wild audio specifically chosen for acoustic difficulty. It uses a “lattice of variations” framework. This allows for multiple valid transcriptions for each utterance. Initial evaluations of 10 state-of-the-art ASR models using Vimarsha revealed substantial shifts in model rankings compared to traditional benchmarks. The benchmark also exposed geographic and demographic performance disparities, along with systematic failure modes across different languages and speech patterns. This development builds on earlier efforts like Indic DiarBench, an AI benchmark for Indian language speech recognition published in August 2026. Vimarsha adds a layer of real-world acoustic diversity and linguistic flexibility to the evaluation process.
Diagnosing Misogyny in Code-Mixed Hinglish
Distinguishing between the use of a slur and its mention (e.g., in counter-speech) poses a hurdle for content moderation. Ashanvi Yadav and Shubham Bhardwaj, in a paper submitted on September 6, 2026, investigate this problem within code-mixed Hinglish. Lexicon-driven misogyny detectors, by design, often fail to make this distinction. This can silence discussions about abuse rather than protecting users.
The researchers identified two evaluation artifacts in a publicly available redacted corpus. First, category-encoding anonymization placeholders in the dataset were found to leak labels, allowing a simple no-learning rule to score 1.000. Even after neutralizing these placeholders, they found that misogynistic and benign comments occupied lexically disjoint registers. This meant that a bag-of-words model could achieve a macro-F1 of approximately 1.00 under random cross-validation, but its performance collapsed under template-disjoint evaluation, where the model sees entirely new linguistic structures. To address this, Yadav and Bhardwaj released Hinglish-MGY-Diag, a deterministic generator that produces a 416-item / 163-minimal-pair contrast-set diagnostic across five linguistically motivated categories. This diagnostic provides a tool for developers to test whether their models can truly differentiate between use and mention in the complex context of Hinglish. This is an important step for developing more effective and equitable moderation systems. This work complements previous discussions around LLM safety, such as the SurakshaEval benchmark published in August 2026. It focuses on a specific, nuanced ethical challenge in a prevalent Indian language context.
Probing Hallucinations in Hinglish Large Language Models
Hallucination detection in large language models (LLMs) is an active area of research, with recent methods using linear classifiers trained on an LLM’s internal activations (hidden states) to detect whether a generated answer is faithful to its input. While these methods have reported 0.90-1.00 AUROC across several benchmarks and languages in 2026, their efficacy with code-mixed input has remained unexplored.
Tanveer Singh, in a paper submitted on August 25, 2026, directly addresses this gap by asking whether a hallucination probe trained on clean-language hidden states can transfer effectively to Hinglish. This is a pertinent question. A large segment of chatbot users in India communicate using Hindi-English code-mixed text. Singh constructed a 5,674-item Hindi/English/Hinglish QA benchmark and generated 17,022 model responses across three open-weight 7-8B LLMs: Qwen2.5-7B, Mistral-7B, and Llama-3.1-8B. The researcher then extracted per-layer hidden states at two token positions and trained linear and MLP probes for both in-distribution detection and cross-lingual transfer. The research examines how well the truthfulness signal within an LLM’s internal representation holds up when confronted with the linguistic blend of Hinglish. This builds on earlier discussions of Indian LLM benchmarks, including those focused on factual knowledge and evaluation maturity, published in August 2026.
These three papers show a maturing research focus within Indian language AI. The work confronts real-world challenges: noisy audio, linguistic ambiguity, and the ethical nuances of code-mixing. The development of Vimarsha, Hinglish-MGY-Diag, and the Hinglish hallucination benchmark provides specific tools and insights for the next generation of AI systems for the region.
Sources
- Vimarsha: Faithful ASR Evaluation for Indian Languages with Demographic Diversity, In-the-Wild Audio and Spelling Variations
- Used, Mentioned, or Condemned? A Controlled Contrast-Set Diagnostic for the Use-Mention Distinction in Code-Mixed Hinglish Misogyny Detection
- Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in Hinglish
Compiled by Launch91 Desk from the sources linked above. More about Launch91.