METAL: Towards Multilingual Meta-Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hada, Rishav, Gumma, Varun, Ahmed, Mohamed, Bali, Kalika, Sitaram, Sunayana |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
di: Hada, Rishav, et al.
Pubblicazione: (2023)
di: Hada, Rishav, et al.
Pubblicazione: (2023)
Contamination Report for Multilingual Benchmarks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2023)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2023)
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
di: Ochieng, Millicent, et al.
Pubblicazione: (2024)
di: Ochieng, Millicent, et al.
Pubblicazione: (2024)
MunTTS: A Text-to-Speech System for Mundari
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
di: Seth, Agrima, et al.
Pubblicazione: (2024)
di: Seth, Agrima, et al.
Pubblicazione: (2024)
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
di: Watts, Ishaan, et al.
Pubblicazione: (2024)
di: Watts, Ishaan, et al.
Pubblicazione: (2024)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
di: Jain, Prachi, et al.
Pubblicazione: (2024)
di: Jain, Prachi, et al.
Pubblicazione: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
di: Seth, Agrima, et al.
Pubblicazione: (2025)
di: Seth, Agrima, et al.
Pubblicazione: (2025)
Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
di: Hada, Rishav, et al.
Pubblicazione: (2024)
di: Hada, Rishav, et al.
Pubblicazione: (2024)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
di: Kumar, Somnath, et al.
Pubblicazione: (2023)
di: Kumar, Somnath, et al.
Pubblicazione: (2023)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
di: Kumar, Somnath, et al.
Pubblicazione: (2024)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
di: Hamna, Hamna, et al.
Pubblicazione: (2025)
di: Hamna, Hamna, et al.
Pubblicazione: (2025)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
di: Mukherjee, Sagnik, et al.
Pubblicazione: (2024)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
di: Schneider, Florian, et al.
Pubblicazione: (2024)
di: Schneider, Florian, et al.
Pubblicazione: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
di: Chitale, Pranjal A., et al.
Pubblicazione: (2025)
di: Chitale, Pranjal A., et al.
Pubblicazione: (2025)
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
di: Avinash, Karthik, et al.
Pubblicazione: (2025)
di: Avinash, Karthik, et al.
Pubblicazione: (2025)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
DEPART: DEcomposing PARiTy across Multilingual LLMs
di: Uppadhyay, Manan, et al.
Pubblicazione: (2026)
di: Uppadhyay, Manan, et al.
Pubblicazione: (2026)
Improving Self Consistency in LLMs through Probabilistic Tokenization
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
di: Sathe, Ashutosh, et al.
Pubblicazione: (2024)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
di: Yadav, Hemant, et al.
Pubblicazione: (2024)
A Multilingual, Culture-First Approach to Addressing Misgendering in LLM Applications
di: Sitaram, Sunayana, et al.
Pubblicazione: (2025)
di: Sitaram, Sunayana, et al.
Pubblicazione: (2025)
Teaching LLMs to Abstain across Languages via Multilingual Feedback
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
di: Feng, Shangbin, et al.
Pubblicazione: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
di: Agarwal, Dhruv, et al.
Pubblicazione: (2025)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
di: Aggarwal, Divyanshu, et al.
Pubblicazione: (2024)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
di: Rajore, Tanmay, et al.
Pubblicazione: (2024)
di: Rajore, Tanmay, et al.
Pubblicazione: (2024)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024)
CultureLLM: Incorporating Cultural Differences into Large Language Models
di: Li, Cheng, et al.
Pubblicazione: (2024)
di: Li, Cheng, et al.
Pubblicazione: (2024)
KAHANI: Culturally-Nuanced Visual Storytelling Tool for Non-Western Cultures
di: Hamna, et al.
Pubblicazione: (2024)
di: Hamna, et al.
Pubblicazione: (2024)
MLMA: Towards Multilingual ASR With Mamba-based Architectures
di: Ali, Mohamed Nabih, et al.
Pubblicazione: (2025)
di: Ali, Mohamed Nabih, et al.
Pubblicazione: (2025)
ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
di: Joshi, Neha, et al.
Pubblicazione: (2025)
di: Joshi, Neha, et al.
Pubblicazione: (2025)
Towards Inclusive NLP: Assessing Compressed Multilingual Transformers across Diverse Language Benchmarks
di: Alshehhi, Maitha, et al.
Pubblicazione: (2025)
di: Alshehhi, Maitha, et al.
Pubblicazione: (2025)
MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge and Reward Models
di: Son, Guijin, et al.
Pubblicazione: (2024)
di: Son, Guijin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
di: Hada, Rishav, et al.
Pubblicazione: (2023) -
Contamination Report for Multilingual Benchmarks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2024) -
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
di: Gumma, Varun, et al.
Pubblicazione: (2024) -
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
di: Ahuja, Sanchit, et al.
Pubblicazione: (2023) -
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
di: Gumma, Varun, et al.
Pubblicazione: (2024)