PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Watts, Ishaan, Gumma, Varun, Yadavalli, Aditya, Seshadri, Vivek, Swaminathan, Manohar, Sitaram, Sunayana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Contamination Report for Multilingual Benchmarks
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
METAL: Towards Multilingual Meta-Evaluation
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
von: Hada, Rishav, et al.
Veröffentlicht: (2023)
von: Hada, Rishav, et al.
Veröffentlicht: (2023)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
von: Jain, Prachi, et al.
Veröffentlicht: (2024)
von: Jain, Prachi, et al.
Veröffentlicht: (2024)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2023)
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2023)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
von: Ochieng, Millicent, et al.
Veröffentlicht: (2024)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
von: Schneider, Florian, et al.
Veröffentlicht: (2024)
von: Schneider, Florian, et al.
Veröffentlicht: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
von: Chitale, Pranjal A., et al.
Veröffentlicht: (2025)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
von: Rajore, Tanmay, et al.
Veröffentlicht: (2024)
von: Rajore, Tanmay, et al.
Veröffentlicht: (2024)
Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
von: Hada, Rishav, et al.
Veröffentlicht: (2024)
CultureLLM: Incorporating Cultural Differences into Large Language Models
von: Li, Cheng, et al.
Veröffentlicht: (2024)
von: Li, Cheng, et al.
Veröffentlicht: (2024)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
von: Sitaram, Sunayana, et al.
Veröffentlicht: (2025)
von: Sitaram, Sunayana, et al.
Veröffentlicht: (2025)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
von: Agarwal, Dhruv, et al.
Veröffentlicht: (2025)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
DEPART: DEcomposing PARiTy across Multilingual LLMs
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
von: Uppadhyay, Manan, et al.
Veröffentlicht: (2026)
Improving Self Consistency in LLMs through Probabilistic Tokenization
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
von: Sathe, Ashutosh, et al.
Veröffentlicht: (2024)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
von: Seth, Agrima, et al.
Veröffentlicht: (2025)
von: Seth, Agrima, et al.
Veröffentlicht: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2024)
von: Mukherjee, Sagnik, et al.
Veröffentlicht: (2024)
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
von: Seth, Agrima, et al.
Veröffentlicht: (2024)
von: Seth, Agrima, et al.
Veröffentlicht: (2024)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
von: Kumar, Somnath, et al.
Veröffentlicht: (2023)
Teaching LLMs to Abstain across Languages via Multilingual Feedback
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
von: Feng, Shangbin, et al.
Veröffentlicht: (2024)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath, et al.
Veröffentlicht: (2024)
Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
von: Joshi, Neha, et al.
Veröffentlicht: (2025)
von: Joshi, Neha, et al.
Veröffentlicht: (2025)
Dynamic Template Selection for Output Token Generation Optimization: MLP-Based and Transformer Approaches
von: Yadavalli, Bharadwaj
Veröffentlicht: (2025)
von: Yadavalli, Bharadwaj
Veröffentlicht: (2025)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
von: Hamna, Hamna, et al.
Veröffentlicht: (2025)
Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
von: Gupta, Ashim, et al.
Veröffentlicht: (2025)
Disentangling Language and Culture for Evaluating Multilingual Large Language Models
von: Ying, Jiahao, et al.
Veröffentlicht: (2025)
von: Ying, Jiahao, et al.
Veröffentlicht: (2025)
ParamBench: A Graduate-Level Benchmark for Evaluating LLM Understanding on Indic Subjects
von: Maheshwari, Ayush, et al.
Veröffentlicht: (2025)
von: Maheshwari, Ayush, et al.
Veröffentlicht: (2025)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024)
Leveraging LLM For Synchronizing Information Across Multilingual Tables
von: Khincha, Siddharth, et al.
Veröffentlicht: (2025)
von: Khincha, Siddharth, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
von: Gumma, Varun, et al.
Veröffentlicht: (2024) -
Contamination Report for Multilingual Benchmarks
von: Ahuja, Sanchit, et al.
Veröffentlicht: (2024) -
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024) -
METAL: Towards Multilingual Meta-Evaluation
von: Hada, Rishav, et al.
Veröffentlicht: (2024) -
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
von: Hada, Rishav, et al.
Veröffentlicht: (2023)