L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages
Fuente:
arXiv
Saved in:
| Main Authors: | Mirashi, Aishwarya, Sonavane, Srushti, Lingayat, Purva, Padhiyar, Tejas, Joshi, Raviraj |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On Importance of Pruning and Distillation for Efficient Low Resource NLP
by: Mirashi, Aishwarya, et al.
Published: (2024)
by: Mirashi, Aishwarya, et al.
Published: (2024)
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
by: Mirashi, Aishwarya, et al.
Published: (2025)
by: Mirashi, Aishwarya, et al.
Published: (2025)
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
by: Rohera, Pritika, et al.
Published: (2024)
by: Rohera, Pritika, et al.
Published: (2024)
L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi
by: Mittal, Saloni, et al.
Published: (2024)
by: Mittal, Saloni, et al.
Published: (2024)
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
by: Endait, Sharvi, et al.
Published: (2025)
by: Endait, Sharvi, et al.
Published: (2025)
L3Cube-IndicHeadline-ID: A Dataset for Headline Identification and Semantic Evaluation in Low-Resource Indian Languages
by: Tanksale, Nishant, et al.
Published: (2025)
by: Tanksale, Nishant, et al.
Published: (2025)
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
L3Cube-MahaEmotions: A Marathi Emotion Recognition Dataset with Synthetic Annotations using CoTR prompting and Large Language Models
by: Kowtal, Nidhi, et al.
Published: (2025)
by: Kowtal, Nidhi, et al.
Published: (2025)
Pralekha: Cross-Lingual Document Alignment for Indic Languages
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
by: Suryanarayanan, Sanjay, et al.
Published: (2024)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
IndicEval-XL: Bridging Linguistic Diversity in Code Generation Across Indic Languages
by: Singh, Ujjwal, et al.
Published: (2025)
by: Singh, Ujjwal, et al.
Published: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
by: Nigam, Shubham Kumar, et al.
Published: (2026)
by: Nigam, Shubham Kumar, et al.
Published: (2026)
L3Cube-MahaSum: A Comprehensive Dataset and BART Models for Abstractive Text Summarization in Marathi
by: Deshmukh, Pranita, et al.
Published: (2024)
by: Deshmukh, Pranita, et al.
Published: (2024)
Long-context Non-factoid Question Answering in Indic Languages
by: Mishra, Ritwik, et al.
Published: (2025)
by: Mishra, Ritwik, et al.
Published: (2025)
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
by: Joshi, Neha, et al.
Published: (2025)
by: Joshi, Neha, et al.
Published: (2025)
Analysis of Indic Language Capabilities in LLMs
by: Vaidya, Aatman, et al.
Published: (2025)
by: Vaidya, Aatman, et al.
Published: (2025)
Statistical Machine Translation for Indic Languages
by: Das, Sudhansu Bala, et al.
Published: (2023)
by: Das, Sudhansu Bala, et al.
Published: (2023)
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages
by: Javed, Tahir, et al.
Published: (2024)
by: Javed, Tahir, et al.
Published: (2024)
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs
by: Khan, Mohammad Aflah, et al.
Published: (2024)
by: Khan, Mohammad Aflah, et al.
Published: (2024)
Navigating Text-to-Image Generative Bias across Indic Languages
by: Mittal, Surbhi, et al.
Published: (2024)
by: Mittal, Surbhi, et al.
Published: (2024)
Safer in Translation? Presupposition Robustness in Indic Languages
by: Palnitkar, Aadi, et al.
Published: (2025)
by: Palnitkar, Aadi, et al.
Published: (2025)
Unicode Normalization and Grapheme Parsing of Indic Languages
by: Ansary, Nazmuddoha, et al.
Published: (2023)
by: Ansary, Nazmuddoha, et al.
Published: (2023)
Overview of the 2023 ICON Shared Task on Gendered Abuse Detection in Indic Languages
by: Vaidya, Aatman, et al.
Published: (2024)
by: Vaidya, Aatman, et al.
Published: (2024)
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
by: Chaudhari, Harsh, et al.
Published: (2023)
by: Chaudhari, Harsh, et al.
Published: (2023)
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages
by: Dawar, Aviral, et al.
Published: (2026)
by: Dawar, Aviral, et al.
Published: (2026)
LittiChoQA: Literary Texts in Indic Languages Chosen for Question Answering
by: Khandelwal, Aarya, et al.
Published: (2026)
by: Khandelwal, Aarya, et al.
Published: (2026)
MMCFND: Multimodal Multilingual Caption-aware Fake News Detection for Low-resource Indic Languages
by: Bansal, Shubhi, et al.
Published: (2024)
by: Bansal, Shubhi, et al.
Published: (2024)
Chain-of-Translation Prompting (CoTR): A Novel Prompting Technique for Low Resource Languages
by: Deshpande, Tejas, et al.
Published: (2024)
by: Deshpande, Tejas, et al.
Published: (2024)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
by: Aravapalli, Akhilesh, et al.
Published: (2024)
by: Aravapalli, Akhilesh, et al.
Published: (2024)
Table Question Answering for Low-resourced Indic Languages
by: Pal, Vaishali, et al.
Published: (2024)
by: Pal, Vaishali, et al.
Published: (2024)
IndicRAGSuite: Large-Scale Datasets and a Benchmark for Indian Language RAG Systems
by: Prasanjith, Pasunuti, et al.
Published: (2025)
by: Prasanjith, Pasunuti, et al.
Published: (2025)
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
by: Menta, Venkata Pushpak Teja
Published: (2026)
by: Menta, Venkata Pushpak Teja
Published: (2026)
Towards Visually-Guided Movie Subtitle Translation for Indic Languages
by: Chintada, Tarun, et al.
Published: (2026)
by: Chintada, Tarun, et al.
Published: (2026)
MILU: A Multi-task Indic Language Understanding Benchmark
by: Verma, Sshubam, et al.
Published: (2024)
by: Verma, Sshubam, et al.
Published: (2024)
Graph-Assisted Culturally Adaptable Idiomatic Translation for Indic Languages
by: Singh, Pratik Rakesh, et al.
Published: (2025)
by: Singh, Pratik Rakesh, et al.
Published: (2025)
Improving Multilingual Neural Machine Translation System for Indic Languages
by: Das, Sudhansu Bala, et al.
Published: (2022)
by: Das, Sudhansu Bala, et al.
Published: (2022)
IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
by: Khan, Mohammed Safi Ur Rahman, et al.
Published: (2024)
Similar Items
-
On Importance of Pruning and Distillation for Efficient Low Resource NLP
by: Mirashi, Aishwarya, et al.
Published: (2024) -
L3Cube-MahaSTS: A Marathi Sentence Similarity Dataset and Models
by: Mirashi, Aishwarya, et al.
Published: (2025) -
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
by: Rohera, Pritika, et al.
Published: (2024) -
L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi
by: Mittal, Saloni, et al.
Published: (2024) -
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
by: Endait, Sharvi, et al.
Published: (2025)