ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
Fuente:
arXiv
Saved in:
| Main Authors: | Maheshwari, Ayush, Sharma, Kaushal, Patel, Vivek, Maheshwari, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025)
by: Maheshwari, Ayush, et al.
Published: (2025)
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026)
by: Maheshwari, Aditya, et al.
Published: (2026)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022)
by: Maheshwari, Ayush, et al.
Published: (2022)
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
by: Roy, Aniruddha, et al.
Published: (2024)
by: Roy, Aniruddha, et al.
Published: (2024)
IndicEval-XL: Bridging Linguistic Diversity in Code Generation Across Indic Languages
by: Singh, Ujjwal, et al.
Published: (2025)
by: Singh, Ujjwal, et al.
Published: (2025)
Sāmayik: A Benchmark and Dataset for English-Sanskrit Translation
by: Maheshwari, Ayush, et al.
Published: (2023)
by: Maheshwari, Ayush, et al.
Published: (2023)
BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains
by: Devane, Vijay, et al.
Published: (2025)
by: Devane, Vijay, et al.
Published: (2025)
ARISE: Iterative Rule Induction and Synthetic Data Generation for Text Classification
by: M., Yashwanth, et al.
Published: (2025)
by: M., Yashwanth, et al.
Published: (2025)
IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languages
by: Jayakumar, Thanmay, et al.
Published: (2026)
by: Jayakumar, Thanmay, et al.
Published: (2026)
Efficacy of Synthetic Data as a Benchmark
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
MILU: A Multi-task Indic Language Understanding Benchmark
by: Verma, Sshubam, et al.
Published: (2024)
by: Verma, Sshubam, et al.
Published: (2024)
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
by: Pattnayak, Priyaranjan, et al.
Published: (2026)
by: Pattnayak, Priyaranjan, et al.
Published: (2026)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
by: KJ, Sankalp, et al.
Published: (2025)
by: KJ, Sankalp, et al.
Published: (2025)
Curate-Train-Refine: A Closed-Loop Agentic Framework for Zero Shot Classification
by: Maheshwari, Gaurav, et al.
Published: (2026)
by: Maheshwari, Gaurav, et al.
Published: (2026)
Collaborating Action by Action: A Multi-agent LLM Framework for Embodied Reasoning
by: White, Isadora, et al.
Published: (2025)
by: White, Isadora, et al.
Published: (2025)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
by: Guo, Jiacheng, et al.
Published: (2025)
by: Guo, Jiacheng, et al.
Published: (2025)
Presentations are not always linear! GNN meets LLM for Document-to-Presentation Transformation with Attribution
by: Maheshwari, Himanshu, et al.
Published: (2024)
by: Maheshwari, Himanshu, et al.
Published: (2024)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
LexGen: Domain-aware Multilingual Lexicon Generation
by: Maheshwari, Ayush, et al.
Published: (2024)
by: Maheshwari, Ayush, et al.
Published: (2024)
IndicSQuAD: A Comprehensive Multilingual Question Answering Dataset for Indic Languages
by: Endait, Sharvi, et al.
Published: (2025)
by: Endait, Sharvi, et al.
Published: (2025)
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
L3Cube-IndicQuest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLMs in Indic Context
by: Rohera, Pritika, et al.
Published: (2024)
by: Rohera, Pritika, et al.
Published: (2024)
ASR Benchmarking: Need for a More Representative Conversational Dataset
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction
by: Maheshwari, Harsh, et al.
Published: (2025)
by: Maheshwari, Harsh, et al.
Published: (2025)
DeliberationBench: When Do More Voices Hurt? A Controlled Study of Multi-LLM Deliberation Protocols
by: Kaushal, Vaarunay, et al.
Published: (2025)
by: Kaushal, Vaarunay, et al.
Published: (2025)
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
by: Joshi, Neha, et al.
Published: (2025)
by: Joshi, Neha, et al.
Published: (2025)
Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
by: Juvekar, Kush, et al.
Published: (2026)
by: Juvekar, Kush, et al.
Published: (2026)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
by: Maurya, Kaushal Kumar, et al.
Published: (2024)
TwinFormer: A Dual-Level Transformer for Long-Sequence Time-Series Forecasting
by: Kumavat, Mahima, et al.
Published: (2025)
by: Kumavat, Mahima, et al.
Published: (2025)
Reliable Answers for Recurring Questions: Boosting Text-to-SQL Accuracy with Template Constrained Decoding
by: Jivani, Smit, et al.
Published: (2026)
by: Jivani, Smit, et al.
Published: (2026)
Social Media Ready Caption Generation for Brands
by: Maheshwari, Himanshu, et al.
Published: (2024)
by: Maheshwari, Himanshu, et al.
Published: (2024)
NC-Bench: An LLM Benchmark for Evaluating Conversational Competence
by: Moore, Robert J., et al.
Published: (2026)
by: Moore, Robert J., et al.
Published: (2026)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
by: Nigam, Shubham Kumar, et al.
Published: (2026)
by: Nigam, Shubham Kumar, et al.
Published: (2026)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
Enhancing Presentation Slide Generation by LLMs with a Multi-Staged End-to-End Approach
by: Bandyopadhyay, Sambaran, et al.
Published: (2024)
by: Bandyopadhyay, Sambaran, et al.
Published: (2024)
Evaluating Cross-lingual Knowledge Consistency in Code-Mixed vis-a-vis Indian Languages using IndicKLAR
by: Mazumder, Debajyoti, et al.
Published: (2026)
by: Mazumder, Debajyoti, et al.
Published: (2026)
Similar Items
-
IndicParam: Benchmark to evaluate LLMs on low-resource Indic Languages
by: Maheshwari, Ayush, et al.
Published: (2025) -
SectEval: Evaluating the Latent Sectarian Preferences of Large Language Models
by: Maheshwari, Aditya, et al.
Published: (2026) -
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024) -
DICTDIS: Dictionary Constrained Disambiguation for Improved NMT
by: Maheshwari, Ayush, et al.
Published: (2022) -
Enhancing Low-Resource NMT with a Multilingual Encoder and Knowledge Distillation: A Case Study
by: Roy, Aniruddha, et al.
Published: (2024)