Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Yanbin, Zhou, Kun, Wang, Zhen, Zhang, Xiangdong, Shao, Yifei, Hao, Shibo, Gu, Yi, Liu, Jieyuan, Singla, Somanshu, Liu, Tianyang, Xing, Eric P., Liu, Zhengzhong, Jin, Haojian, Hu, Zhiting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
by: Singla, Somanshu, et al.
Published: (2024)
by: Singla, Somanshu, et al.
Published: (2024)
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
by: Hao, Shibo, et al.
Published: (2023)
by: Hao, Shibo, et al.
Published: (2023)
World Reasoning Arena
by: PAN Team, et al.
Published: (2026)
by: PAN Team, et al.
Published: (2026)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
by: Zha, Yuheng, et al.
Published: (2025)
by: Zha, Yuheng, et al.
Published: (2025)
Pandora: Towards General World Model with Natural Language Actions and Video States
by: Xiang, Jiannan, et al.
Published: (2024)
by: Xiang, Jiannan, et al.
Published: (2024)
Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings
by: Shao, Yifei, et al.
Published: (2026)
by: Shao, Yifei, et al.
Published: (2026)
Data Mixing Optimization for Supervised Fine-Tuning of Large Language Models
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
scPilot: Large Language Model Reasoning Toward Automated Single-Cell Analysis and Discovery
by: Gao, Yiming, et al.
Published: (2026)
by: Gao, Yiming, et al.
Published: (2026)
Towards General Continuous Memory for Vision-Language Models
by: Wu, Wenyi, et al.
Published: (2025)
by: Wu, Wenyi, et al.
Published: (2025)
Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
by: Zhao, Zekai, et al.
Published: (2025)
by: Zhao, Zekai, et al.
Published: (2025)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
FIRE-Bench: Evaluating Agents on the Rediscovery of Scientific Insights
by: Wang, Zhen, et al.
Published: (2026)
by: Wang, Zhen, et al.
Published: (2026)
PlanetServe: A Decentralized, Scalable, and Privacy-Preserving Overlay for Democratizing Large Language Model Serving
by: Fang, Fei, et al.
Published: (2025)
by: Fang, Fei, et al.
Published: (2025)
Environmental Uncertainty Modulates the Framing Effect in Planning: Evidence From Staged Neural Correlates
by: Zhican He, et al.
Published: (2025)
by: Zhican He, et al.
Published: (2025)
How Does Controllability Emerge In Language Models During Pretraining?
by: She, Jianshu, et al.
Published: (2025)
by: She, Jianshu, et al.
Published: (2025)
Democratic Transformation and the Vernacular Public Arena in India
Published: (2022)
Published: (2022)
Democratic Transformation and the Vernacular Public Arena in India
Published: (2022)
Published: (2022)
Driving Strategy Using an Improved Ant Colony System for Energy‐Efficient Train
by: Chengda Yang, et al.
Published: (2024)
by: Chengda Yang, et al.
Published: (2024)
CellMaster: Collaborative Cell Type Annotation in Single-Cell Analysis
by: Wang, Zhen, et al.
Published: (2026)
by: Wang, Zhen, et al.
Published: (2026)
VIRENA: Virtual Arena for Research, Education, and Democratic Innovation
by: Hoes, Emma, et al.
Published: (2026)
by: Hoes, Emma, et al.
Published: (2026)
ANVIL: Accelerator-Native Video Interpolation via Codec Motion Vector Priors
by: Liu, Shibo
Published: (2026)
by: Liu, Shibo
Published: (2026)
Multiple solutions for Schr{ö}dinger-Poisson-Slater equations with critical growth
by: Liu, Shibo
Published: (2025)
by: Liu, Shibo
Published: (2025)
On the integral formula of the Jacobian determinant
by: Liu, Shibo
Published: (2025)
by: Liu, Shibo
Published: (2025)
Physical derivation of the coarea formula and an elementary proof via gradient flow
by: Liu, Shibo
Published: (2026)
by: Liu, Shibo
Published: (2026)
Astragalus: Automatic Configuration Repair for Production Networks
by: Gu, Zhenrong, et al.
Published: (2026)
by: Gu, Zhenrong, et al.
Published: (2026)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
by: Cheng, Zhoujun, et al.
Published: (2026)
by: Cheng, Zhoujun, et al.
Published: (2026)
Effect of Argon Flow Rate on Power Consumption of a 120‐t Ladle Furnace
by: Jian Song, et al.
Published: (2024)
by: Jian Song, et al.
Published: (2024)
Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective
by: Cheng, Zhoujun, et al.
Published: (2025)
by: Cheng, Zhoujun, et al.
Published: (2025)
A Generative Model Enhanced Multi-Agent Reinforcement Learning Method for Electric Vehicle Charging Navigation
by: Qi, Tianyang, et al.
Published: (2025)
by: Qi, Tianyang, et al.
Published: (2025)
HumanStudy-Bench: Towards AI Agent Design for Participant Simulation
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
Critiques of World Models
by: Xing, Eric, et al.
Published: (2025)
by: Xing, Eric, et al.
Published: (2025)
General Agentic Planning Through Simulative Reasoning with World Models
by: Deng, Mingkai, et al.
Published: (2025)
by: Deng, Mingkai, et al.
Published: (2025)
FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding via Agent-of-Thoughts Reasoning
by: Chen, Haodong, et al.
Published: (2025)
by: Chen, Haodong, et al.
Published: (2025)
CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments
by: Liu, Xuan, et al.
Published: (2025)
by: Liu, Xuan, et al.
Published: (2025)
When Agent Markets Arrive
by: Liu, Xuan, et al.
Published: (2026)
by: Liu, Xuan, et al.
Published: (2026)
Lattica: A Decentralized Cross-NAT Communication Framework for Scalable AI Inference and Training
by: Yang, Ween, et al.
Published: (2025)
by: Yang, Ween, et al.
Published: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
by: Liu, Qin, et al.
Published: (2025)
by: Liu, Qin, et al.
Published: (2025)
GameArena: Evaluating LLM Reasoning through Live Computer Games
by: Hu, Lanxiang, et al.
Published: (2024)
by: Hu, Lanxiang, et al.
Published: (2024)
Democratizing High-Fidelity Co-Speech Gesture Video Generation
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
Similar Items
-
Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models
by: Singla, Somanshu, et al.
Published: (2024) -
ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings
by: Hao, Shibo, et al.
Published: (2023) -
World Reasoning Arena
by: PAN Team, et al.
Published: (2026) -
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
by: Zha, Yuheng, et al.
Published: (2025) -
Pandora: Towards General World Model with Natural Language Actions and Video States
by: Xiang, Jiannan, et al.
Published: (2024)