Large language models (LLMs) often reproduce social biases, a problem especially critical in India where demographic factors like caste and urban-rural location significantly influence economic outcomes. A new benchmark, RupeeBias, has been introduced to audit demographic bias in LLM-generated economic guidance, according to research submitted on September 25, 2026, by Pavithra P M Nair, Bhavik Talaviya, Shourya Bhushan, Rahul Pankajakshan, Seema Guruvadoo, Avinash Agarwal, Gilad Gressel, and Krishnashree Achuthan. This work addresses a specific, nuanced aspect of AI safety, extending previous efforts like the SurakshaEval benchmark for general Indian language LLM safety, which was covered in August.

Existing LLM bias benchmarks are largely designed around Western demographic categories, missing key axes of economic disparity relevant to India. RupeeBias directly targets this gap. Biased economic advice from LLMs could influence users’ perceptions of their own worth in financial negotiations, from loan comparisons to salary requests. RupeeBias shows Indian AI research is moving beyond general performance metrics, now focusing on ethical implications and societal impact within India’s diverse population.

Benchmarking Full-Duplex Voice Agents

IndicFDB offers a new benchmark for evaluating full-duplex voice agents across ten Indian languages. Submitted on September 25, 2026, by Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh, Manmeet Kaur, Sagar Jain, and Hanuman Sidh, this work extends the English-only Full-Duplex-Bench, which struggled with the linguistic and technical requirements of Indian languages.

Full-duplex voice agents need to handle pauses, manage turn-taking, provide backchannels, and respond to user interruptions in real time. The original Full-Duplex-Bench relied on word-timestamped Automatic Speech Recognition (ASR) and an English-prompted LLM judge, making it unsuitable for multilingual applications. IndicFDB, first reported on arXiv, identifies conversational events in multilingual speech, evaluates timing without reliable word-level alignment, and judges responses across diverse languages. The benchmark features 12,350 samples, nearly 17 times the size of the original Full-Duplex-Bench. To construct this, the authors mined approximately 50,000 hours of channel-separated conversations using voice activity detection (VAD) and created human-validated synthetic user interruption samples. This advancement is critical for developing more natural and responsive voice assistants capable of fluid interaction in Indian languages, similar in ambition to earlier speech benchmarks like Indic DiarBench but focused on the complexities of real-time dialogue.

New Word Similarity Datasets for Six Indian Languages

Foundational evaluation tools for natural language processing also continue to expand with the introduction of new word similarity datasets for six Indian languages. Researchers Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee, Arjit Srivastava, and M. Shrivastava presented manually annotated monolingual word similarity datasets for Urdu, Telugu, Marathi, Punjabi, Tamil, and Gujarati. These languages are among the most spoken Indian languages globally, after Hindi and Bengali. The paper, submitted on September 28, 2026, and available on arXiv, addresses a long-standing need for reliable evaluation metrics for word representations.

The quality of word representations, or embeddings, is important for numerous downstream NLP tasks, and word similarity tasks are a popular method for evaluating them. The authors constructed these datasets by translating and re-annotating existing English word similarity datasets. They also provided baseline scores for state-of-the-art word representation models, evaluated on these newly created datasets for Urdu, Telugu, and Marathi. This work directly supports the development and refinement of word embeddings specific to these languages, which can improve the performance of everything from search engines to machine translation systems operating in the Indian linguistic context.

These three distinct contributions show a concentrated effort within Indian AI research to build out foundational infrastructure for Indian language AI. This includes specialized benchmarks for ethical auditing, detailed datasets for complex conversational agents, and fundamental evaluation tools for core NLP components. The RupeeBias benchmark, for instance, specifically targets demographic categories like caste and urban-rural location. These are often overlooked in Western-centric models but are critical for fair economic guidance in India.

Compiled by Launch91 Desk from the sources linked above. More about Launch91.