The assessment of India’s foundation models now extends to the maturity of their evaluation methods, alongside model capabilities. Recent research introduces new benchmarks for factual knowledge, offers detailed case studies on prompt injection safety in Indian language models, and proposes a framework for understanding how well models are even being measured. These efforts collectively point to a growing emphasis on systematic and critical evaluation in the development of Indian language AI.
L3Cube-IndicQuest v2: Measuring India-Specific Factual Knowledge
Researchers Rinit Jain, Tirthraj Mahajan, Advait Joshi, and Raviraj Joshi have released L3Cube-IndicQuest v2, a new benchmark for evaluating the India-specific factual knowledge of large language models (LLMs). The benchmark, detailed in a paper submitted on August 16, 2026, comprises 3,471 English question-answer pairs. These pairs cover nine domains, drawn from educational curricula, competitive examination materials, and domain-specific reference books. The team used a hybrid construction method involving LLM-based question generation and validation, semantic deduplication, and human verification to ensure quality and scalability.
L3Cube-IndicQuest v2 is translated into 19 Indic languages, resulting in a publicly available multilingual dataset of 69,420 question-answer pairs across 20 languages. This scale allows for a broad evaluation of LLMs’ ability to process and recall information relevant to India in multiple regional languages. The authors evaluated six LLMs using three protocols, including LLM-as-a-judge and two deterministic lexical methods, providing a baseline for future comparisons.
Sarvam-105B Examined for Indirect Prompt Injection
A separate case study by Madhusudhanan G, submitted on August 15, 2026, investigated the reliability of chain-of-thought monitoring as a safety signal for indirect prompt injection. The study focused on Sarvam-105B across English, Tamil, and Tanglish (code-mixed Tamil and English). This research, conducted with eight manually verified synthetic scenarios, one model, one annotator, and a deterministic generation seed, provides specific, if limited, data on how visible reasoning affects prompt injection success.
A pilot phase with four scenarios found 5 out of 12 injected attack successes occurred without visible reasoning, and 1 out of 11 occurred with reasoning. A preregistered follow-up phase, also with four scenarios, showed a reversed trend: 2 out of 12 attacks succeeded without reasoning, while 3 out of 12 succeeded with reasoning. The author notes that with only four scenarios per phase, this design cannot distinguish a real reasoning-mode effect from prompt-specific variation or sampling noise. However, an observation across 20 non-empty injected-thinking traces revealed that all 17 benign-correct outputs stated an intent to ignore the injection, while all three attack successes stated an intent to follow it. This specific finding suggests a potential signal for monitoring, despite the study’s small scale. This kind of granular safety research complements broader efforts like the SurakshaEval benchmark, which this publication covered on August 13, 2026, by looking deeply into a single model’s behavior under specific attack vectors.
Assessing the Evaluation Maturity of Indian Foundation Models
Avinash Agarwal and Vridhi Jain, in their paper submitted on August 12, 2026, present a framework that shifts focus from merely assessing model capabilities to also assessing the maturity of the evaluation processes themselves. Their work, titled “Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework,” addresses how governments fund indigenous foundation models for national AI capability and multilingual computing.
The paper offers a structured, benchmark-based comparative assessment of Indian foundation models against global frontier and comparable-scale models across eight capability domains: general-purpose reasoning, coding and software engineering, agentic AI and computer use, cybersecurity, vision and image understanding, video and multimodal understanding, scientific research, and Indic language capability. Using only publicly reported results, the authors propose an exploratory four-dimension Benchmark Maturity Index (BMI). This index scores each domain on standardization, participation, and independent verification, highlighting that apparent capability gaps might also reflect gaps in evaluation maturity. This perspective shows that the utility of benchmarks like L3Cube-IndicQuest v2 and the insights from safety studies on models like Sarvam-105B depend heavily on the strength and transparency of their application and reporting.
Sources
- L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages
- Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
- Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Compiled by Launch91 Desk from the sources linked above. More about Launch91.