ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Qin, Dineen, Jacob, Huang, Yuxi, Zhang, Sheng, Poon, Hoifung, Zhou, Ben, Chen, Muhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
di: Zhou, Wenxuan, et al.
Pubblicazione: (2023)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2023)
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
di: Huang, James Y., et al.
Pubblicazione: (2025)
di: Huang, James Y., et al.
Pubblicazione: (2025)
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
di: Xu, Nan, et al.
Pubblicazione: (2024)
di: Xu, Nan, et al.
Pubblicazione: (2024)
Offset Unlearning for Large Language Models
di: Huang, James Y., et al.
Pubblicazione: (2024)
di: Huang, James Y., et al.
Pubblicazione: (2024)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
di: Li, Bangzheng, et al.
Pubblicazione: (2025)
MetaScale: Test-Time Scaling with Evolving Meta-Thoughts
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
di: Huang, James Y., et al.
Pubblicazione: (2025)
di: Huang, James Y., et al.
Pubblicazione: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
AutoBencher: Towards Declarative Benchmark Construction
di: Li, Xiang Lisa, et al.
Pubblicazione: (2024)
di: Li, Xiang Lisa, et al.
Pubblicazione: (2024)
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
di: Dineen, Jacob, et al.
Pubblicazione: (2026)
di: Dineen, Jacob, et al.
Pubblicazione: (2026)
Exploring Scaling Laws for EHR Foundation Models
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
di: Zhang, Sheng, et al.
Pubblicazione: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
di: Dineen, Jacob, et al.
Pubblicazione: (2025)
di: Dineen, Jacob, et al.
Pubblicazione: (2025)
DocLens: Multi-aspect Fine-grained Evaluation for Medical Text Generation
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
di: Xie, Yiqing, et al.
Pubblicazione: (2023)
Evaluating Medical LLMs by Levels of Autonomy: A Survey Moving from Benchmarks to Applications
di: Ye, Xiao, et al.
Pubblicazione: (2025)
di: Ye, Xiao, et al.
Pubblicazione: (2025)
Two Heads are Better than One: Nested PoE for Robust Defense Against Multi-Backdoors
di: Graf, Victoria, et al.
Pubblicazione: (2024)
di: Graf, Victoria, et al.
Pubblicazione: (2024)
Pareto Optimal Learning for Estimating Large Language Model Errors
di: Zhao, Theodore, et al.
Pubblicazione: (2023)
di: Zhao, Theodore, et al.
Pubblicazione: (2023)
BOW: Reinforcement Learning for Bottlenecked Next Word Prediction
di: Shen, Ming, et al.
Pubblicazione: (2025)
di: Shen, Ming, et al.
Pubblicazione: (2025)
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
di: Yin, Yanbin, et al.
Pubblicazione: (2025)
di: Yin, Yanbin, et al.
Pubblicazione: (2025)
Familiarity-Aware Evidence Compression for Retrieval-Augmented Generation
di: Jung, Dongwon, et al.
Pubblicazione: (2024)
di: Jung, Dongwon, et al.
Pubblicazione: (2024)
T-Rex: Text-assisted Retrosynthesis Prediction
di: Liu, Yifeng, et al.
Pubblicazione: (2024)
di: Liu, Yifeng, et al.
Pubblicazione: (2024)
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
di: Tong, Terry, et al.
Pubblicazione: (2024)
di: Tong, Terry, et al.
Pubblicazione: (2024)
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
di: Xu, Wendong, et al.
Pubblicazione: (2025)
di: Xu, Wendong, et al.
Pubblicazione: (2025)
MathArena: Evaluating LLMs on Uncontaminated Math Competitions
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
di: Balunović, Mislav, et al.
Pubblicazione: (2025)
Attribute Structuring Improves LLM-Based Evaluation of Clinical Text Summaries
di: Gero, Zelalem, et al.
Pubblicazione: (2024)
di: Gero, Zelalem, et al.
Pubblicazione: (2024)
MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
di: Du, Yuhao, et al.
Pubblicazione: (2025)
di: Du, Yuhao, et al.
Pubblicazione: (2025)
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
di: Zhu, Yakun, et al.
Pubblicazione: (2025)
di: Zhu, Yakun, et al.
Pubblicazione: (2025)
OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI
di: Huang, Zhen, et al.
Pubblicazione: (2024)
di: Huang, Zhen, et al.
Pubblicazione: (2024)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
di: Wang, Cheng, et al.
Pubblicazione: (2026)
di: Wang, Cheng, et al.
Pubblicazione: (2026)
OdysseyArena: Benchmarking Large Language Models For Long-Horizon, Active and Inductive Interactions
di: Xu, Fangzhi, et al.
Pubblicazione: (2026)
di: Xu, Fangzhi, et al.
Pubblicazione: (2026)
Video Models Can Reason with Verifiable Rewards
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
di: Zhu, Tinghui, et al.
Pubblicazione: (2026)
RECAP: Transparent Inference-Time Emotion Alignment for Medical Dialogue Systems
di: Srinivasan, Adarsh, et al.
Pubblicazione: (2025)
di: Srinivasan, Adarsh, et al.
Pubblicazione: (2025)
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
di: Thakur, Nandan, et al.
Pubblicazione: (2024)
di: Thakur, Nandan, et al.
Pubblicazione: (2024)
ToW: Thoughts of Words Improve Reasoning in Large Language Models
di: Xu, Zhikun, et al.
Pubblicazione: (2024)
di: Xu, Zhikun, et al.
Pubblicazione: (2024)
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
di: Wang, Cheng, et al.
Pubblicazione: (2025)
di: Wang, Cheng, et al.
Pubblicazione: (2025)
Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
di: Dekoninck, Jasper, et al.
Pubblicazione: (2026)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
di: Chen, Junjie, et al.
Pubblicazione: (2026)
di: Chen, Junjie, et al.
Pubblicazione: (2026)
SudoLM: Learning Access Control of Parametric Knowledge with Authorization Alignment
di: Liu, Qin, et al.
Pubblicazione: (2024)
di: Liu, Qin, et al.
Pubblicazione: (2024)
Bencher: Simple and Reproducible Benchmarking for Black-Box Optimization
di: Papenmeier, Leonard, et al.
Pubblicazione: (2025)
di: Papenmeier, Leonard, et al.
Pubblicazione: (2025)
Documenti analoghi
-
UniversalNER: Targeted Distillation from Large Language Models for Open Named Entity Recognition
di: Zhou, Wenxuan, et al.
Pubblicazione: (2023) -
Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
di: Huang, James Y., et al.
Pubblicazione: (2025) -
From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
di: Xu, Nan, et al.
Pubblicazione: (2024) -
Offset Unlearning for Large Language Models
di: Huang, James Y., et al.
Pubblicazione: (2024) -
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
di: Li, Bangzheng, et al.
Pubblicazione: (2025)