ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qin, Dineen, Jacob, Huang, Yuxi, Zhang, Sheng, Poon, Hoifung, Zhou, Ben, Chen, Muhao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024)
von: Xu, Nan, et al.
Veröffentlicht: (2024)
Offset Unlearning for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024)
von: Huang, James Y., et al.
Veröffentlicht: (2024)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
von: Liu, Qin, et al.
Veröffentlicht: (2025)
von: Liu, Qin, et al.
Veröffentlicht: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
von: Huang, James Y., et al.
Veröffentlicht: (2025)
von: Huang, James Y., et al.
Veröffentlicht: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
AutoBencher: Towards Declarative Benchmark Construction
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2024)
von: Li, Xiang Lisa, et al.
Veröffentlicht: (2024)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
von: Dineen, Jacob, et al.
Veröffentlicht: (2026)
von: Dineen, Jacob, et al.
Veröffentlicht: (2026)
Exploring Scaling Laws for EHR Foundation Models
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
von: Zhang, Sheng, et al.
Veröffentlicht: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
von: Dineen, Jacob, et al.
Veröffentlicht: (2025)
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
von: Xie, Yiqing, et al.
Veröffentlicht: (2023)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
von: Ye, Xiao, et al.
Veröffentlicht: (2025)
von: Ye, Xiao, et al.
Veröffentlicht: (2025)
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
von: Graf, Victoria, et al.
Veröffentlicht: (2024)
von: Graf, Victoria, et al.
Veröffentlicht: (2024)
Pareto Optimal Learning for Estimating Large Language Model Errors
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
von: Zhao, Theodore, et al.
Veröffentlicht: (2023)
BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
von: Shen, Ming, et al.
Veröffentlicht: (2025)
von: Shen, Ming, et al.
Veröffentlicht: (2025)
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
von: Yin, Yanbin, et al.
Veröffentlicht: (2025)
von: Yin, Yanbin, et al.
Veröffentlicht: (2025)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
von: Jung, Dongwon, et al.
Veröffentlicht: (2024)
von: Jung, Dongwon, et al.
Veröffentlicht: (2024)
T-Rex: Text-assisted Retrosynthesis Prediction
von: Liu, Yifeng, et al.
Veröffentlicht: (2024)
von: Liu, Yifeng, et al.
Veröffentlicht: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
von: Tong, Terry, et al.
Veröffentlicht: (2024)
von: Tong, Terry, et al.
Veröffentlicht: (2024)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
von: Xu, Wendong, et al.
Veröffentlicht: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
von: Balunović, Mislav, et al.
Veröffentlicht: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
von: Gero, Zelalem, et al.
Veröffentlicht: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
von: Zhu, Yakun, et al.
Veröffentlicht: (2025)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
von: Huang, Zhen, et al.
Veröffentlicht: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
von: Xu, Fangzhi, et al.
Veröffentlicht: (2026)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2026)
Video Models Can Reason with Verifiable Rewards
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
von: Srinivasan, Adarsh, et al.
Veröffentlicht: (2025)
von: Srinivasan, Adarsh, et al.
Veröffentlicht: (2025)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
von: Thakur, Nandan, et al.
Veröffentlicht: (2024)
ToW: Thoughts of Words Improve Reasoning in Large Language Models
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
von: Xu, Zhikun, et al.
Veröffentlicht: (2024)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
von: Wang, Cheng, et al.
Veröffentlicht: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
von: Dekoninck, Jasper, et al.
Veröffentlicht: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
von: Liu, Qin, et al.
Veröffentlicht: (2024)
von: Liu, Qin, et al.
Veröffentlicht: (2024)
Bencher: Simple and Reproducible Benchmarking for Black-Box Optimization
von: Papenmeier, Leonard, et al.
Veröffentlicht: (2025)
von: Papenmeier, Leonard, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2023) -
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
von: Huang, James Y., et al.
Veröffentlicht: (2025) -
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
von: Xu, Nan, et al.
Veröffentlicht: (2024) -
Offset Unlearning for Large Language Models
von: Huang, James Y., et al.
Veröffentlicht: (2024) -
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)