BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Lingfeng, Lu, Yunlong, Zhang, Yuefei, Yao, Jingyu, Zhu, Yixin, Cheng, KeYuan, Wang, Yongyi, Zheng, Qirui, Yang, Xionghui, Li, Wenxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
Decoupling Return-to-Go for Efficient Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
Style-Preserving Policy Optimization for Game Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
di: Wang, Yongyi, et al.
Pubblicazione: (2026)
Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
di: Li, Lingfeng, et al.
Pubblicazione: (2025)
Pareto-guided Pipeline for Distilling Featherweight AI Agents in Mobile MOBA Games
di: Yang, Xionghui, et al.
Pubblicazione: (2026)
di: Yang, Xionghui, et al.
Pubblicazione: (2026)
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
di: Zheng, Qirui, et al.
Pubblicazione: (2025)
di: Zheng, Qirui, et al.
Pubblicazione: (2025)
Adapting Rules of Official International Mahjong for Online Players
di: Wang, Chucai, et al.
Pubblicazione: (2026)
di: Wang, Chucai, et al.
Pubblicazione: (2026)
Constructing Non-Markovian Decision Process via History Aggregator
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
di: Wang, Yongyi, et al.
Pubblicazione: (2025)
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
di: Bu, Yuyan, et al.
Pubblicazione: (2026)
AutoMonitor-Bench: Evaluating the Reliability of LLM-Based Misbehavior Monitor
di: Yang, Shu, et al.
Pubblicazione: (2026)
di: Yang, Shu, et al.
Pubblicazione: (2026)
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
di: Chen, Yuefei, et al.
Pubblicazione: (2025)
di: Chen, Yuefei, et al.
Pubblicazione: (2025)
A General Anchor-Based Framework for Scalable Fair Clustering
di: Wei, Shengfei, et al.
Pubblicazione: (2025)
di: Wei, Shengfei, et al.
Pubblicazione: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
di: Zhang, Yuzhe, et al.
Pubblicazione: (2026)
di: Zhang, Yuzhe, et al.
Pubblicazione: (2026)
Confidence Intervals for Evaluation of Data Mining
di: Yuan, Zheng, et al.
Pubblicazione: (2025)
di: Yuan, Zheng, et al.
Pubblicazione: (2025)
ShuttleEnv: An Interactive Data-Driven RL Environment for Badminton Strategy Modeling
di: Li, Ang, et al.
Pubblicazione: (2026)
di: Li, Ang, et al.
Pubblicazione: (2026)
GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation
di: Li, Baiqi, et al.
Pubblicazione: (2024)
di: Li, Baiqi, et al.
Pubblicazione: (2024)
AnchorDrive: LLM Scenario Rollout with Anchor-Guided Diffusion Regeneration for Safety-Critical Scenario Generation
di: Jiang, Zhulin, et al.
Pubblicazione: (2026)
di: Jiang, Zhulin, et al.
Pubblicazione: (2026)
LifelongAgentBench: Evaluating LLM Agents as Lifelong Learners
di: Zheng, Junhao, et al.
Pubblicazione: (2025)
di: Zheng, Junhao, et al.
Pubblicazione: (2025)
Secret Stealing Attacks on Local LLM Fine-Tuning through Supply-Chain Model Code Backdoors
di: Li, Zi, et al.
Pubblicazione: (2026)
di: Li, Zi, et al.
Pubblicazione: (2026)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
Efficient HDR Reconstruction from Real-World Raw Images
di: Yang, Qirui, et al.
Pubblicazione: (2023)
di: Yang, Qirui, et al.
Pubblicazione: (2023)
HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
di: Sturgeon, Benjamin, et al.
Pubblicazione: (2025)
di: Sturgeon, Benjamin, et al.
Pubblicazione: (2025)
A Proof of the Biquadratic Linear AFL for GL(4)
di: Li, Qirui
Pubblicazione: (2025)
di: Li, Qirui
Pubblicazione: (2025)
When and What to Ask: AskBench and Rubric-Guided RLVR for LLM Clarification
di: Zhao, Jiale, et al.
Pubblicazione: (2026)
di: Zhao, Jiale, et al.
Pubblicazione: (2026)
AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
di: Li, Zheng, et al.
Pubblicazione: (2025)
di: Li, Zheng, et al.
Pubblicazione: (2025)
A Review of the Bearing Characteristics and Failure Mechanisms of Anchor Plates and Group Anchor Systems
di: Hongyang Huang, et al.
Pubblicazione: (2025)
di: Hongyang Huang, et al.
Pubblicazione: (2025)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
di: Cheng, Xiang, et al.
Pubblicazione: (2026)
di: Cheng, Xiang, et al.
Pubblicazione: (2026)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
di: Costarelli, Anthony, et al.
Pubblicazione: (2024)
di: Costarelli, Anthony, et al.
Pubblicazione: (2024)
Cross-Cultural Communication in the Digital Age: An Analysis of Cultural Representation and Inclusivity in Emojis
di: Li, Lingfeng, et al.
Pubblicazione: (2024)
di: Li, Lingfeng, et al.
Pubblicazione: (2024)
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
di: Hao, Zhuoyuan, et al.
Pubblicazione: (2026)
di: Hao, Zhuoyuan, et al.
Pubblicazione: (2026)
Solving Motion Planning Tasks with a Scalable Generative Model
di: Hu, Yihan, et al.
Pubblicazione: (2024)
di: Hu, Yihan, et al.
Pubblicazione: (2024)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
di: Li, Haitao, et al.
Pubblicazione: (2024)
di: Li, Haitao, et al.
Pubblicazione: (2024)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
di: Ge, Mingji, et al.
Pubblicazione: (2026)
di: Ge, Mingji, et al.
Pubblicazione: (2026)
RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
di: Liu, Yantao, et al.
Pubblicazione: (2024)
di: Liu, Yantao, et al.
Pubblicazione: (2024)
Deep Ideation: Designing LLM Agents to Generate Novel Research Ideas on Scientific Concept Network
di: Zhao, Keyu, et al.
Pubblicazione: (2025)
di: Zhao, Keyu, et al.
Pubblicazione: (2025)
Systematically Evaluating Cell‐Free DNA Fragmentation Patterns for Cancer Diagnosis and Enhanced Cancer Detection via Integrating Multiple Fragmentation Patterns
di: Yuying Hou, et al.
Pubblicazione: (2024)
di: Yuying Hou, et al.
Pubblicazione: (2024)
FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems
di: Xing, Shanli, et al.
Pubblicazione: (2026)
di: Xing, Shanli, et al.
Pubblicazione: (2026)
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
di: Guo, Zhengkang, et al.
Pubblicazione: (2026)
Correction to “Color‐Tunable Fluorescent Hierarchical Nanoassemblies with Concentration‐Encoded Emission”
di: Jia Kong, et al.
Pubblicazione: (2024)
di: Jia Kong, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Synthetic POMDPs to Challenge Memory-Augmented RL: Memory Demand Structure Modeling
di: Wang, Yongyi, et al.
Pubblicazione: (2025) -
Decoupling Return-to-Go for Efficient Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026) -
Style-Preserving Policy Optimization for Game Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025) -
Beyond Autoregressive RTG: Conditioning via Injection Outside Sequential Modeling in Decision Transformer
di: Wang, Yongyi, et al.
Pubblicazione: (2026) -
Mxplainer: Explain and Learn Insights by Imitating Mahjong Agents
di: Li, Lingfeng, et al.
Pubblicazione: (2025)