MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ni, Jinjie, Xue, Fuzhao, Yue, Xiang, Deng, Yuntian, Shah, Mahir, Jain, Kabir, Neubig, Graham, You, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
Boosting LLM via Learning from Data Iteratively and Selectively
von: Jia, Qi, et al.
Veröffentlicht: (2024)
von: Jia, Qi, et al.
Veröffentlicht: (2024)
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024)
The Wisdom of Partisan Crowds: Comparing Collective Intelligence in Humans and LLM-based Agents
von: Chuang, Yun-Shiuan, et al.
Veröffentlicht: (2023)
von: Chuang, Yun-Shiuan, et al.
Veröffentlicht: (2023)
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
von: Li, Yisen, et al.
Veröffentlicht: (2025)
von: Li, Yisen, et al.
Veröffentlicht: (2025)
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd
von: Liu, Shang, et al.
Veröffentlicht: (2024)
von: Liu, Shang, et al.
Veröffentlicht: (2024)
Mix-of-Language-Experts Architecture for Multilingual Programming
von: Zong, Yifan, et al.
Veröffentlicht: (2025)
von: Zong, Yifan, et al.
Veröffentlicht: (2025)
AcademicEval: Live Long-Context LLM Benchmark
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2025)
Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds
von: Abels, Axel, et al.
Veröffentlicht: (2025)
von: Abels, Axel, et al.
Veröffentlicht: (2025)
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2023)
von: Plaza-del-Arco, Flor Miriam, et al.
Veröffentlicht: (2023)
FedEval-LLM: Federated Evaluation of Large Language Models on Downstream Tasks with Collective Wisdom
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
von: He, Yuanqin, et al.
Veröffentlicht: (2024)
Go-Browse: Training Web Agents with Structured Exploration
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
von: Gandhi, Apurva, et al.
Veröffentlicht: (2025)
BehaviorBox: Automated Discovery of Fine-Grained Performance Differences Between Language Models
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2025)
von: Tjuatja, Lindia, et al.
Veröffentlicht: (2025)
ClusterFusion: Hybrid Clustering with Embedding Guidance and LLM Adaptation
von: Xu, Yiming, et al.
Veröffentlicht: (2025)
von: Xu, Yiming, et al.
Veröffentlicht: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
von: Yeo, Edward, et al.
Veröffentlicht: (2025)
Wisdom of the Crowds in Forecasting: Forecast Summarization for Supporting Future Event Prediction
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
von: Saha, Anisha, et al.
Veröffentlicht: (2025)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Effective Strategies for Asynchronous Software Engineering Agents
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
von: Geng, Jiayi, et al.
Veröffentlicht: (2026)
TokDrift: When LLM Speaks in Subwords but Code Speaks in Grammar
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
von: Li, Yinxi, et al.
Veröffentlicht: (2025)
DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
von: Liang, Chumeng, et al.
Veröffentlicht: (2025)
Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
von: Huan, Maggie, et al.
Veröffentlicht: (2025)
StackEval: Benchmarking LLMs in Coding Assistance
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
von: Shah, Nidhish, et al.
Veröffentlicht: (2024)
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
GaelEval: Benchmarking LLM Performance for Scottish Gaelic
von: Devine, Peter, et al.
Veröffentlicht: (2026)
von: Devine, Peter, et al.
Veröffentlicht: (2026)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
von: Xia, Chunqiu Steven, et al.
Veröffentlicht: (2024)
Low-Rank Adapters Meet Neural Architecture Search for LLM Compression
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
von: Muñoz, J. Pablo, et al.
Veröffentlicht: (2025)
Benchmarking Failures in Tool-Augmented Language Models
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
von: Treviño, Eduardo, et al.
Veröffentlicht: (2025)
An Incomplete Loop: Instruction Inference, Instruction Following, and In-context Learning in Language Models
von: Liu, Emmy, et al.
Veröffentlicht: (2024)
von: Liu, Emmy, et al.
Veröffentlicht: (2024)
What Is Missing in Multilingual Visual Reasoning and How to Fix It
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
von: Song, Yueqi, et al.
Veröffentlicht: (2024)
Midtraining Bridges Pretraining and Posttraining Distributions
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
von: Liu, Emmy, et al.
Veröffentlicht: (2025)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2023)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2023)
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
von: Onohara, Shota, et al.
Veröffentlicht: (2024)
DependEval: Benchmarking LLMs for Repository Dependency Understanding
von: Du, Junjia, et al.
Veröffentlicht: (2025)
von: Du, Junjia, et al.
Veröffentlicht: (2025)
Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2025)
OneEval: Benchmarking LLM Knowledge-intensive Reasoning over Diverse Knowledge Bases
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
von: Chen, Yongrui, et al.
Veröffentlicht: (2025)
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
von: Yue, Xiang, et al.
Veröffentlicht: (2024)
SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
von: Petrukha, Ivan, et al.
Veröffentlicht: (2025)
Training Versatile Coding Agents in Synthetic Environments
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
von: Zhu, Yiqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024) -
Boosting LLM via Learning from Data Iteratively and Selectively
von: Jia, Qi, et al.
Veröffentlicht: (2024) -
OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
von: Xue, Fuzhao, et al.
Veröffentlicht: (2024) -
The Wisdom of Partisan Crowds: Comparing Collective Intelligence in Humans and LLM-based Agents
von: Chuang, Yun-Shiuan, et al.
Veröffentlicht: (2023) -
On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
von: Zhang, Charlie, et al.
Veröffentlicht: (2025)