A recent replication study found that a semantic-extractive summarization method, when adapted for Hindi, performed significantly worse than a simple three-sentence lead baseline. This finding, reported by Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, and Niyaz Ahmad Wani on September 24, 2026, shows the difficulty in adapting existing natural language processing techniques to Indian languages. Rigorous evaluation is key.
The researchers replicated the distributional-semantics extractive summarization method of Mohd, Jan and Shah (2020), substituting Devanagari-appropriate components for each language-specific step. They evaluated the adapted system on two independent corpora: the Hindi portion of XL-Sum and FIRE ILSUM 2.0 Hindi. Their evaluation used a Devanagari-aware ROUGE implementation, which they validated against the XL-Sum authors’ own multilingual scorer. All comparisons were drawn using 1000-resample paired bootstraps.
The replicated system, in its published equal-weight configuration, trailed the Lead-3 baseline by 0.042 ROUGE-1 F-score on XL-Sum and by a larger 0.265 ROUGE-1 F-score on ILSUM. A feature ablation study, where individual components of the summarization method were removed, showed that sentence position was the only feature that contributed to the system’s performance. The researchers found that using sentence position alone could exactly reproduce the lead baseline’s performance. This suggests that the more complex semantic-extractive components did not provide a measurable benefit over simply extracting the first three sentences for Hindi. The full paper is available on arXiv.
This replication study shows a broader issue in Indian language AI. It needs strong evaluation and foundational data resources that capture linguistic and cultural nuances. Complex models require careful verification against simpler baselines and diverse datasets to prove their worth in real-world Indic language contexts.
Kshetrimayum Boynao Singh, Nitin Kumar Mishra, Palash Pratim Dutta, Atai Waris Khan, Aparna Kaushik, Avinash Kumar, Deeksha, and Deepak Kumar presented COILD. This new Indic-centric parallel corpus and benchmark for machine translation (MT) addresses limitations of existing multilingual resources. Existing resources are largely built from English-pivot content and often fail to capture the diversity of Indian languages. Their work was submitted on September 23, 2026.
COILD comprises over 1.16 million human-translated and human-verified sentence pairs. This corpus covers 20 Indian language pairs, spanning the Indo-Aryan, Dravidian, Tibeto-Burman, and Austro-Asiatic language families. The researchers emphasized that COILD is built entirely from original Indian language sources, collected from licensed repositories across eight distinct domains chosen for their direct real-world applicability. This approach aims to create a resource that better represents Indian languages linguistically and culturally, moving beyond English-centric data collection methods. The COILD paper is available on arXiv.
COILD provides a critical resource for training and evaluating machine translation models. It offers a large dataset, tailored to the complexities of inter-Indian language translation. For developers building AI systems for languages like Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia, and Assamese, resources like COILD are essential. The dataset focuses on direct Indian language pairs, not relying on English as an intermediary. This could lead to more accurate and contextually appropriate translations, especially for applications needing translation between two Indian languages without an English pivot.
These two papers illustrate the ongoing, fundamental work required in Indian language AI research. The replication study shows the value of rigorous testing, even for established methods, to ensure they provide actual benefits over simpler alternatives for Indic scripts and languages. Meanwhile, the COILD corpus meets the need for high-quality, domain-specific, and culturally relevant data to build advanced models. The 1.16 million sentence pairs in COILD provide a concrete step towards more effective machine translation for India’s many languages.
Compiled by Launch91 Desk from the sources linked above. More about Launch91.