MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
Fuente:
arXiv
Guardado en:
| Autores principales: | Ahuja, Sanchit, Aggarwal, Divyanshu, Gumma, Varun, Watts, Ishaan, Sathe, Ashutosh, Ochieng, Millicent, Hada, Rishav, Jain, Prachi, Axmed, Maxamed, Bali, Kalika, Sitaram, Sunayana |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
por: Jain, Prachi, et al.
Publicado: (2024)
por: Jain, Prachi, et al.
Publicado: (2024)
METAL: Towards Multilingual Meta-Evaluation
por: Hada, Rishav, et al.
Publicado: (2024)
por: Hada, Rishav, et al.
Publicado: (2024)
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
Contamination Report for Multilingual Benchmarks
por: Ahuja, Sanchit, et al.
Publicado: (2024)
por: Ahuja, Sanchit, et al.
Publicado: (2024)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
por: Sathe, Ashutosh, et al.
Publicado: (2024)
por: Sathe, Ashutosh, et al.
Publicado: (2024)
Are Large Language Model-based Evaluators the Solution to Scaling Up Multilingual Evaluation?
por: Hada, Rishav, et al.
Publicado: (2023)
por: Hada, Rishav, et al.
Publicado: (2023)
DOSA: A Dataset of Social Artifacts from Different Indian Geographical Subcultures
por: Seth, Agrima, et al.
Publicado: (2024)
por: Seth, Agrima, et al.
Publicado: (2024)
Improving Self Consistency in LLMs through Probabilistic Tokenization
por: Sathe, Ashutosh, et al.
Publicado: (2024)
por: Sathe, Ashutosh, et al.
Publicado: (2024)
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios
por: Ochieng, Millicent, et al.
Publicado: (2024)
por: Ochieng, Millicent, et al.
Publicado: (2024)
HEALTH-PARIKSHA: Assessing RAG Models for Health Chatbots in Real-World Multilingual Settings
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languages
por: Chitale, Pranjal A., et al.
Publicado: (2025)
por: Chitale, Pranjal A., et al.
Publicado: (2025)
PARIKSHA: A Large-Scale Investigation of Human-LLM Evaluator Agreement on Multilingual and Multi-Cultural Data
por: Watts, Ishaan, et al.
Publicado: (2024)
por: Watts, Ishaan, et al.
Publicado: (2024)
Towards Inducing Long-Context Abilities in Multilingual Neural Machine Translation Models
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
MunTTS: A Text-to-Speech System for Mundari
por: Gumma, Varun, et al.
Publicado: (2024)
por: Gumma, Varun, et al.
Publicado: (2024)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
por: Aggarwal, Divyanshu, et al.
Publicado: (2024)
Bridging the Language Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
por: Kumar, Somnath, et al.
Publicado: (2023)
por: Kumar, Somnath, et al.
Publicado: (2023)
Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology
por: Hada, Rishav, et al.
Publicado: (2024)
por: Hada, Rishav, et al.
Publicado: (2024)
Bridging the Gap: Dynamic Learning Strategies for Improving Multilingual Performance in LLMs
por: Kumar, Somnath, et al.
Publicado: (2024)
por: Kumar, Somnath, et al.
Publicado: (2024)
How Deep Is Representational Bias in LLMs? The Cases of Caste and Religion
por: Seth, Agrima, et al.
Publicado: (2025)
por: Seth, Agrima, et al.
Publicado: (2025)
M5 -- A Diverse Benchmark to Assess the Performance of Large Multimodal Models Across Multilingual and Multicultural Vision-Language Tasks
por: Schneider, Florian, et al.
Publicado: (2024)
por: Schneider, Florian, et al.
Publicado: (2024)
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting
por: Mukherjee, Sagnik, et al.
Publicado: (2024)
por: Mukherjee, Sagnik, et al.
Publicado: (2024)
Prompt Engineering a Prompt Engineer
por: Ye, Qinyuan, et al.
Publicado: (2023)
por: Ye, Qinyuan, et al.
Publicado: (2023)
Parameter Alignment Mitigates Catastrophic Forgetting in Multilingual Expert Language Models
por: Ahuja, Sanchit, et al.
Publicado: (2026)
por: Ahuja, Sanchit, et al.
Publicado: (2026)
Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settings
por: Hamna, Hamna, et al.
Publicado: (2025)
por: Hamna, Hamna, et al.
Publicado: (2025)
Efficient Training of Language Models with Compact and Consistent Next Token Distributions
por: Sathe, Ashutosh, et al.
Publicado: (2024)
por: Sathe, Ashutosh, et al.
Publicado: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
por: Yadav, Hemant, et al.
Publicado: (2024)
por: Yadav, Hemant, et al.
Publicado: (2024)
Geometry of Decision Making in Language Models
por: Joshi, Abhinav, et al.
Publicado: (2025)
por: Joshi, Abhinav, et al.
Publicado: (2025)
CultureLLM: Incorporating Cultural Differences into Large Language Models
por: Li, Cheng, et al.
Publicado: (2024)
por: Li, Cheng, et al.
Publicado: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
por: Yadav, Hemant, et al.
Publicado: (2023)
por: Yadav, Hemant, et al.
Publicado: (2023)
Can Small Language Models Use What They Retrieve? An Empirical Study of Retrieval Utilization Across Model Scale
por: Pandey, Sanchit
Publicado: (2026)
por: Pandey, Sanchit
Publicado: (2026)
sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
por: Ahuja, Sanchit, et al.
Publicado: (2024)
por: Ahuja, Sanchit, et al.
Publicado: (2024)
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
por: Yadav, Hemant, et al.
Publicado: (2024)
por: Yadav, Hemant, et al.
Publicado: (2024)
JOOCI: a Framework for Learning Comprehensive Speech Representations
por: Yadav, Hemant, et al.
Publicado: (2024)
por: Yadav, Hemant, et al.
Publicado: (2024)
Protect: Towards Robust Guardrailing Stack for Trustworthy Enterprise LLM Systems
por: Avinash, Karthik, et al.
Publicado: (2025)
por: Avinash, Karthik, et al.
Publicado: (2025)
OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!
por: Lei, Jingdi, et al.
Publicado: (2025)
por: Lei, Jingdi, et al.
Publicado: (2025)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
por: Kumar, Divyanshu, et al.
Publicado: (2024)
por: Kumar, Divyanshu, et al.
Publicado: (2024)
Scaling Laws for Multilingual Language Models
por: He, Yifei, et al.
Publicado: (2024)
por: He, Yifei, et al.
Publicado: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
por: Agarwal, Dhruv, et al.
Publicado: (2025)
por: Agarwal, Dhruv, et al.
Publicado: (2025)
Study on Gravitational Waves from Binary Mergers and Constraints on the Hubble Parameter
por: Laloo, Rishav, et al.
Publicado: (2025)
por: Laloo, Rishav, et al.
Publicado: (2025)
Ejemplares similares
-
MAPLE: Multilingual Evaluation of Parameter Efficient Finetuning of Large Language Models
por: Aggarwal, Divyanshu, et al.
Publicado: (2024) -
MAFIA: Multi-Adapter Fused Inclusive LanguAge Models
por: Jain, Prachi, et al.
Publicado: (2024) -
METAL: Towards Multilingual Meta-Evaluation
por: Hada, Rishav, et al.
Publicado: (2024) -
Exploring Pretraining via Active Forgetting for Improving Cross Lingual Transfer for Decoder Language Models
por: Aggarwal, Divyanshu, et al.
Publicado: (2024) -
Contamination Report for Multilingual Benchmarks
por: Ahuja, Sanchit, et al.
Publicado: (2024)