The Leaderboard Illusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Shivalika, Nan, Yiyang, Wang, Alex, D'Souza, Daniel, Kapoor, Sayash, Üstün, Ahmet, Koyejo, Sanmi, Deng, Yuntian, Longpre, Shayne, Smith, Noah A., Ermis, Beyza, Fadaee, Marzieh, Hooker, Sara |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SimMerge: Learning to Select Merge Operators from Similarity Signals
von: Bolton, Oliver, et al.
Veröffentlicht: (2026)
von: Bolton, Oliver, et al.
Veröffentlicht: (2026)
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
von: Aakanksha, et al.
Veröffentlicht: (2024)
von: Aakanksha, et al.
Veröffentlicht: (2024)
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
von: Aakanksha, et al.
Veröffentlicht: (2024)
von: Aakanksha, et al.
Veröffentlicht: (2024)
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025)
The Multilingual Divide and Its Impact on Global AI Safety
von: Peppin, Aidan, et al.
Veröffentlicht: (2025)
von: Peppin, Aidan, et al.
Veröffentlicht: (2025)
Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
von: Ahmadian, Arash, et al.
Veröffentlicht: (2024)
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
From One to Many: Expanding the Scope of Toxicity Mitigation in Language Models
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
von: Pozzobon, Luiza, et al.
Veröffentlicht: (2024)
The 2024 Foundation Model Transparency Index
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
von: Üstün, Ahmet, et al.
Veröffentlicht: (2024)
von: Üstün, Ahmet, et al.
Veröffentlicht: (2024)
To Code, or Not To Code? Exploring Impact of Code in Pre-training
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
von: Aryabumi, Viraat, et al.
Veröffentlicht: (2024)
Foundation Model Transparency Reports
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
The 2025 Foundation Model Transparency Index
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
Aya Vision: Advancing the Frontier of Multilingual Multimodality
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
von: Dash, Saurabh, et al.
Veröffentlicht: (2025)
LLM See, LLM Do: Guiding Data Generation to Target Non-Differentiable Objectives
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
von: Shimabucoro, Luísa, et al.
Veröffentlicht: (2024)
Automatic WordNet Construction Using Markov Chain Monte Carlo
von: Marzieh Fadaee
Veröffentlicht: (2013)
von: Marzieh Fadaee
Veröffentlicht: (2013)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
von: Oderinwale, Hamidah, et al.
Veröffentlicht: (2024)
von: Oderinwale, Hamidah, et al.
Veröffentlicht: (2024)
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
von: Abagyan, Diana, et al.
Veröffentlicht: (2025)
von: Abagyan, Diana, et al.
Veröffentlicht: (2025)
Multilingual Arbitrage: Optimizing Data Pools to Accelerate Multilingual Progress
von: Odumakinde, Ayomide, et al.
Veröffentlicht: (2024)
von: Odumakinde, Ayomide, et al.
Veröffentlicht: (2024)
The Reality of AI and Biorisk
von: Peppin, Aidan, et al.
Veröffentlicht: (2024)
von: Peppin, Aidan, et al.
Veröffentlicht: (2024)
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
von: Singh, Shivalika, et al.
Veröffentlicht: (2024)
von: Singh, Shivalika, et al.
Veröffentlicht: (2024)
Verification Limits Code LLM Training
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
von: Gureja, Srishti, et al.
Veröffentlicht: (2025)
The Disparate Impacts of Speculative Decoding
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
von: Sandler, Jameson, et al.
Veröffentlicht: (2025)
Nexus: Specialization meets Adaptability for Efficiently Training Mixture of Experts
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
von: Gritsch, Nikolas, et al.
Veröffentlicht: (2024)
The Art of Asking: Multilingual Prompt Optimization for Synthetic Data
von: Mora, David, et al.
Veröffentlicht: (2025)
von: Mora, David, et al.
Veröffentlicht: (2025)
Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers
von: D'souza, Daniel, et al.
Veröffentlicht: (2025)
von: D'souza, Daniel, et al.
Veröffentlicht: (2025)
Causally Inspired Regularization Enables Domain General Representations
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
von: Salaudeen, Olawale, et al.
Veröffentlicht: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
von: Robertson, Zachary, et al.
Veröffentlicht: (2025)
von: Robertson, Zachary, et al.
Veröffentlicht: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
von: Vo, Truong, et al.
Veröffentlicht: (2025)
von: Vo, Truong, et al.
Veröffentlicht: (2025)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
von: Dang, John, et al.
Veröffentlicht: (2024)
von: Dang, John, et al.
Veröffentlicht: (2024)
Exploring the Latest LLMs for Leaderboard Extraction
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
von: Yu, Simon, et al.
Veröffentlicht: (2024)
von: Yu, Simon, et al.
Veröffentlicht: (2024)
The Limits of Inference Scaling Through Resampling
von: Stroebl, Benedikt, et al.
Veröffentlicht: (2024)
von: Stroebl, Benedikt, et al.
Veröffentlicht: (2024)
Build Agent Advocates, Not Platform Agents
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2025)
Promises and pitfalls of artificial intelligence for legal applications
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
von: Kapoor, Sayash, et al.
Veröffentlicht: (2024)
A Framework for Objective-Driven Dynamical Stochastic Fields
von: Zhang, Yibo Jacky, et al.
Veröffentlicht: (2025)
von: Zhang, Yibo Jacky, et al.
Veröffentlicht: (2025)
Instruction Finetuning for Leaderboard Generation from Empirical AI Research
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
von: Kabongo, Salomon, et al.
Veröffentlicht: (2024)
A Systematic Review of NeurIPS Dataset Management Practices
von: Wu, Yiwei, et al.
Veröffentlicht: (2024)
von: Wu, Yiwei, et al.
Veröffentlicht: (2024)
How Does Quantization Affect Multilingual LLMs?
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
von: Marchisio, Kelly, et al.
Veröffentlicht: (2024)
Tiny Aya: Bridging Scale and Multilingual Depth
von: Salamanca, Alejandro R., et al.
Veröffentlicht: (2026)
von: Salamanca, Alejandro R., et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SimMerge: Learning to Select Merge Operators from Similarity Signals
von: Bolton, Oliver, et al.
Veröffentlicht: (2026) -
Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning
von: Aakanksha, et al.
Veröffentlicht: (2024) -
The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm
von: Aakanksha, et al.
Veröffentlicht: (2024) -
A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics
von: Shimabucoro, Luisa, et al.
Veröffentlicht: (2025) -
The Multilingual Divide and Its Impact on Global AI Safety
von: Peppin, Aidan, et al.
Veröffentlicht: (2025)