BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Lingfeng, Lu, Yunlong, Zhang, Yuefei, Yao, Jingyu, Zhu, Yixin, Cheng, KeYuan, Wang, Yongyi, Zheng, Qirui, Yang, Xionghui, Li, Wenxin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
por: Wang, Yongyi, et al.
Publicado: (2025)
por: Wang, Yongyi, et al.
Publicado: (2025)
Decoupling Return-to-Go for Efficient Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026)
por: Wang, Yongyi, et al.
Publicado: (2026)
Style-Preserving Policy Optimization for Game Agents
por: Li, Lingfeng, et al.
Publicado: (2025)
por: Li, Lingfeng, et al.
Publicado: (2025)
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026)
por: Wang, Yongyi, et al.
Publicado: (2026)
Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
por: Li, Lingfeng, et al.
Publicado: (2025)
por: Li, Lingfeng, et al.
Publicado: (2025)
Pareto-guided Pipeline for Distilling Featherweight AI Agents in Mobile MOBA Games
por: Yang, Xionghui, et al.
Publicado: (2026)
por: Yang, Xionghui, et al.
Publicado: (2026)
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
por: Zheng, Qirui, et al.
Publicado: (2025)
por: Zheng, Qirui, et al.
Publicado: (2025)
Adapting Rules of Official International Mahjong for Online Players
por: Wang, Chucai, et al.
Publicado: (2026)
por: Wang, Chucai, et al.
Publicado: (2026)
Constructing Non-Markovian Decision Process via History Aggregator
por: Wang, Yongyi, et al.
Publicado: (2025)
por: Wang, Yongyi, et al.
Publicado: (2025)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
por: Bu, Yuyan, et al.
Publicado: (2026)
por: Bu, Yuyan, et al.
Publicado: (2026)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
por: Yang, Shu, et al.
Publicado: (2026)
por: Yang, Shu, et al.
Publicado: (2026)
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
por: Chen, Yuefei, et al.
Publicado: (2025)
por: Chen, Yuefei, et al.
Publicado: (2025)
A General Anchor-Based Framework for Scalable Fair Clustering
por: Wei, Shengfei, et al.
Publicado: (2025)
por: Wei, Shengfei, et al.
Publicado: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
por: Zhang, Yuzhe, et al.
Publicado: (2026)
por: Zhang, Yuzhe, et al.
Publicado: (2026)
Confidence Intervals for Evaluation of Data Mining
por: Yuan, Zheng, et al.
Publicado: (2025)
por: Yuan, Zheng, et al.
Publicado: (2025)
ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling
por: Li, Ang, et al.
Publicado: (2026)
por: Li, Ang, et al.
Publicado: (2026)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
por: Li, Baiqi, et al.
Publicado: (2024)
por: Li, Baiqi, et al.
Publicado: (2024)
AnchorDrive: LLM Scenario Rollout with Anchor-Guided Diffusion Regeneration for Safety-Critical Scenario Generation
por: Jiang, Zhulin, et al.
Publicado: (2026)
por: Jiang, Zhulin, et al.
Publicado: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
por: Zheng, Junhao, et al.
Publicado: (2025)
por: Zheng, Junhao, et al.
Publicado: (2025)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
por: Li, Zi, et al.
Publicado: (2026)
por: Li, Zi, et al.
Publicado: (2026)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
por: Markowitz, Elan, et al.
Publicado: (2025)
por: Markowitz, Elan, et al.
Publicado: (2025)
Efficient HDR Reconstruction from Real-World Raw Images
por: Yang, Qirui, et al.
Publicado: (2023)
por: Yang, Qirui, et al.
Publicado: (2023)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
por: Sturgeon, Benjamin, et al.
Publicado: (2025)
por: Sturgeon, Benjamin, et al.
Publicado: (2025)
A Proof of the Biquadratic Linear AFL for GL(4)
por: Li, Qirui
Publicado: (2025)
por: Li, Qirui
Publicado: (2025)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
por: Zhao, Jiale, et al.
Publicado: (2026)
por: Zhao, Jiale, et al.
Publicado: (2026)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
por: Li, Zheng, et al.
Publicado: (2025)
por: Li, Zheng, et al.
Publicado: (2025)
A Review of the Bearing Characteristics and Failure Mechanisms of Anchor Plates and Group Anchor Systems
por: Hongyang Huang, et al.
Publicado: (2025)
por: Hongyang Huang, et al.
Publicado: (2025)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
por: Cheng, Xiang, et al.
Publicado: (2026)
por: Cheng, Xiang, et al.
Publicado: (2026)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
por: Costarelli, Anthony, et al.
Publicado: (2024)
por: Costarelli, Anthony, et al.
Publicado: (2024)
Cross-Cultural Communication in the Digital Age: An Analysis of Cultural Representation and Inclusivity in Emojis
por: Li, Lingfeng, et al.
Publicado: (2024)
por: Li, Lingfeng, et al.
Publicado: (2024)
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
por: Hao, Zhuoyuan, et al.
Publicado: (2026)
por: Hao, Zhuoyuan, et al.
Publicado: (2026)
Solving Motion Planning Tasks with a Scalable Generative Model
por: Hu, Yihan, et al.
Publicado: (2024)
por: Hu, Yihan, et al.
Publicado: (2024)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
por: Li, Haitao, et al.
Publicado: (2024)
por: Li, Haitao, et al.
Publicado: (2024)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
por: Ge, Mingji, et al.
Publicado: (2026)
por: Ge, Mingji, et al.
Publicado: (2026)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
por: Liu, Yantao, et al.
Publicado: (2024)
por: Liu, Yantao, et al.
Publicado: (2024)
Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
por: Zhao, Keyu, et al.
Publicado: (2025)
por: Zhao, Keyu, et al.
Publicado: (2025)
Systematically Evaluating Cell‐Free DNA Fragmentation Patterns for Cancer Diagnosis and Enhanced Cancer Detection via Integrating Multiple Fragmentation Patterns
por: Yuying Hou, et al.
Publicado: (2024)
por: Yuying Hou, et al.
Publicado: (2024)
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
por: Xing, Shanli, et al.
Publicado: (2026)
por: Xing, Shanli, et al.
Publicado: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
por: Guo, Zhengkang, et al.
Publicado: (2026)
por: Guo, Zhengkang, et al.
Publicado: (2026)
Correction to “Color‐Tunable Fluorescent Hierarchical Nanoassemblies with Concentration‐Encoded Emission”
por: Jia Kong, et al.
Publicado: (2024)
por: Jia Kong, et al.
Publicado: (2024)
Ejemplares similares
-
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
por: Wang, Yongyi, et al.
Publicado: (2025) -
Decoupling Return-to-Go for Efficient Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026) -
Style-Preserving Policy Optimization for Game Agents
por: Li, Lingfeng, et al.
Publicado: (2025) -
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
por: Wang, Yongyi, et al.
Publicado: (2026) -
Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
por: Li, Lingfeng, et al.
Publicado: (2025)