Existing safety evaluations for large language models (LLMs) often fall short when applied to the linguistic and cultural nuances of India. A new dataset, SurakshaEval, aims to close this gap by providing a benchmark explicitly designed for ten major Indian languages alongside English. Debopriyo Banerjee, Kapil Rajesh Kavitha, Angana Borah, Xudong Han, Yuxia Wang, Parameswari Krishnamurthy, Utkarsh Agarwal, and Atharva Kulkarni developed this benchmark, as first reported on August 8, 2026, on arXiv.

The challenge with evaluating LLM safety extends beyond simple translation. What is considered harmless in one cultural context may be offensive or even dangerous in another. India, with its vast linguistic and cultural diversity, presents a unique set of safety considerations that English-centric benchmarks typically overlook. This oversight means that models trained and evaluated predominantly on Western data may exhibit significant safety failures when deployed for Indian language users, potentially generating inappropriate, biased, or harmful content.

SurakshaEval: A Culturally Grounded Approach to Safety

SurakshaEval addresses this by compiling human-written prompts that cover real-world scenarios. The benchmark includes prompts for Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, in addition to English. This comprehensive linguistic coverage is crucial for models aiming to serve India’s diverse population.

The researchers designed SurakshaEval with two types of prompts: generic prompts that are broadly applicable across India, and region- and language-specific prompts. The latter category is particularly important because it captures localized sociocultural sensitivities that might not be apparent in a pan-Indian or global context. For instance, a prompt related to dietary restrictions or religious customs might carry different safety implications depending on the specific language and cultural group. This distinction acknowledges that even within India, safety risks are not monolithic.

The authors used SurakshaEval to benchmark a broad range of state-of-the-art LLMs, establishing a baseline for safety performance across these languages. Their work identifies recurring failure modes in these models, indicating that current LLMs struggle to consistently navigate the complex safety challenges of Indian languages. While the specific numerical scores and detailed failure analyses are not available in the summary, the identification of these “recurring failures” points to a systemic issue. It suggests that current approaches to multilingual safety, often relying on direct translation or generalized filters, are insufficient for the specificities of Indian linguistic and cultural contexts.

Why a Dedicated Benchmark Matters for Indian AI

For developers building Indian-language AI, a benchmark like SurakshaEval is fundamental. Without a standardized, culturally aware evaluation tool, it is difficult to measure progress, compare models, or ensure that deployed applications are safe and ethical for their intended users. Relying on benchmarks designed for other linguistic contexts risks deploying models that might generate misinformation, perpetuate stereotypes, or otherwise cause harm in Indian communities.

This work complements the broader push within India to develop indigenous foundation models, a trend driven by governmental initiatives aimed at strengthening national AI capability and digital sovereignty, as noted in a separate arXiv paper by Avinash Agarwal and Vridhi Jain. While that paper points out that Indian models show strong scores on established benchmarks like MMLU and MATH-500, those benchmarks primarily assess general reasoning and coding, not nuanced cultural safety. SurakshaEval directly fills a critical gap for multilingual computing, ensuring that as Indian LLMs advance in capability, they also mature in their ethical and safety considerations.

The development of SurakshaEval also speaks to the challenges highlighted in other recent research, such as Sahil Pardasani and Madhusudan Singh’s work on verifying LLM benchmarks. Their paper points out the potential for bias and unverified claims in model evaluation, even noting how unverified claims about model performance once contributed to significant market panic. By creating a human-written, expert-designed dataset, SurakshaEval reduces reliance on “LLM-as-a-judge” methods, which can introduce identity-aware bias, scoring answers based on source model rather than quality. This human-centric approach to prompt creation lends more credibility to the safety evaluations performed using SurakshaEval.

Implications for LLM Development in India

The findings from SurakshaEval imply that developers and researchers working on Indian-language LLMs need to move beyond generic safety filters. The presence of region- and language-specific failure modes suggests that safety mechanisms must be deeply integrated into the model’s understanding of cultural context, rather than applied as an afterthought. This could involve more granular fine-tuning on diverse, culturally sensitive datasets, or developing language- and region-specific safety policies.

Ultimately, SurakshaEval provides a crucial tool for the responsible development of AI in India. By bringing a structured, culturally informed lens to safety evaluation, it enables researchers and companies to build more reliable, trustworthy, and socially aware LLMs that genuinely serve the needs of India’s diverse linguistic communities. The recurring failures it identifies offer clear directions for future research and development, guiding efforts towards building AI that is not only capable but also safe and equitable.

Compiled by Launch91 Desk from the sources linked above. More about Launch91.