BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Jiahao, Cheng, Fei, Jiang, Junfeng, Yu, Zefan, Aizawa, Akiko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility
por: Huang, Jiahao, et al.
Publicado: (2026)
por: Huang, Jiahao, et al.
Publicado: (2026)
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
por: Jiang, Junfeng, et al.
Publicado: (2024)
por: Jiang, Junfeng, et al.
Publicado: (2024)
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
por: Wu, Yunze, et al.
Publicado: (2025)
por: Wu, Yunze, et al.
Publicado: (2025)
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
por: Costarelli, Anthony, et al.
Publicado: (2024)
por: Costarelli, Anthony, et al.
Publicado: (2024)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
por: Wang, Ruipeng, et al.
Publicado: (2026)
por: Wang, Ruipeng, et al.
Publicado: (2026)
DCA-Bench: A Benchmark for Dataset Curation Agents
por: Huang, Benhao, et al.
Publicado: (2024)
por: Huang, Benhao, et al.
Publicado: (2024)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
por: Shang, Yu, et al.
Publicado: (2025)
por: Shang, Yu, et al.
Publicado: (2025)
Reasoning Depth and Environment Complexity: A Controlled Study of RLVR Data Allocation across Logical Reasoning Tasks
por: Zhu, Yihua, et al.
Publicado: (2026)
por: Zhu, Yihua, et al.
Publicado: (2026)
TPS-Bench: Evaluating AI Agents' Tool Planning \& Scheduling Abilities in Compounding Tasks
por: Xu, Hanwen, et al.
Publicado: (2025)
por: Xu, Hanwen, et al.
Publicado: (2025)
Harnessing PDF Data for Improving Japanese Large Multimodal Models
por: Baek, Jeonghun, et al.
Publicado: (2025)
por: Baek, Jeonghun, et al.
Publicado: (2025)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
por: Zhou, Yifan, et al.
Publicado: (2026)
por: Zhou, Yifan, et al.
Publicado: (2026)
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
por: Xu, Zelai, et al.
Publicado: (2025)
por: Xu, Zelai, et al.
Publicado: (2025)
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
por: Zheng, Tianshi, et al.
Publicado: (2025)
por: Zheng, Tianshi, et al.
Publicado: (2025)
COMMA: A Communicative Multimodal Multi-Agent Benchmark
por: Ossowski, Timothy, et al.
Publicado: (2024)
por: Ossowski, Timothy, et al.
Publicado: (2024)
Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Values
por: Dong, Haonan, et al.
Publicado: (2026)
por: Dong, Haonan, et al.
Publicado: (2026)
TAME: A Trustworthy Test-Time Evolution of Agent Memory with Systematic Benchmarking
por: Cheng, Yu, et al.
Publicado: (2026)
por: Cheng, Yu, et al.
Publicado: (2026)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
por: Liu, Wenrui, et al.
Publicado: (2025)
por: Liu, Wenrui, et al.
Publicado: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
por: Mishra, Prakamya, et al.
Publicado: (2025)
por: Mishra, Prakamya, et al.
Publicado: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
por: Zhang, Hanrong, et al.
Publicado: (2024)
por: Zhang, Hanrong, et al.
Publicado: (2024)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
por: Yin, Sheng, et al.
Publicado: (2024)
por: Yin, Sheng, et al.
Publicado: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
por: Tu, Xinming, et al.
Publicado: (2026)
por: Tu, Xinming, et al.
Publicado: (2026)
T3DM: Test-Time Training-Guided Distribution Shift Modelling for Temporal Knowledge Graph Reasoning
por: Si, Yuehang, et al.
Publicado: (2025)
por: Si, Yuehang, et al.
Publicado: (2025)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
por: Nguyen, Bang, et al.
Publicado: (2026)
por: Nguyen, Bang, et al.
Publicado: (2026)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
por: Jing, Lucas, et al.
Publicado: (2026)
por: Jing, Lucas, et al.
Publicado: (2026)
AutoPenBench: Benchmarking Generative Agents for Penetration Testing
por: Gioacchini, Luca, et al.
Publicado: (2024)
por: Gioacchini, Luca, et al.
Publicado: (2024)
GeoAgentBench: A Dynamic Execution Benchmark for Tool-Augmented Agents in Spatial Analysis
por: Yu, Bo, et al.
Publicado: (2026)
por: Yu, Bo, et al.
Publicado: (2026)
WritingBench: A Comprehensive Benchmark for Generative Writing
por: Wu, Yuning, et al.
Publicado: (2025)
por: Wu, Yuning, et al.
Publicado: (2025)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
por: Kim, Kon Woo, et al.
Publicado: (2025)
por: Kim, Kon Woo, et al.
Publicado: (2025)
SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes
por: Li, Kuan, et al.
Publicado: (2026)
por: Li, Kuan, et al.
Publicado: (2026)
ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents
por: Lee, Seunghyun, et al.
Publicado: (2026)
por: Lee, Seunghyun, et al.
Publicado: (2026)
JFTA-Bench: Evaluate LLM's Ability of Tracking and Analyzing Malfunctions Using Fault Trees
por: Wang, Yuhui, et al.
Publicado: (2026)
por: Wang, Yuhui, et al.
Publicado: (2026)
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
por: Zheng, Qiaoyuan, et al.
Publicado: (2026)
por: Zheng, Qiaoyuan, et al.
Publicado: (2026)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024)
por: Deng, Shihan, et al.
Publicado: (2024)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
por: Jiang, Yixing, et al.
Publicado: (2025)
por: Jiang, Yixing, et al.
Publicado: (2025)
Eigenpruning: an Interpretability-Inspired PEFT Method
por: Vergara-Browne, Tomás, et al.
Publicado: (2024)
por: Vergara-Browne, Tomás, et al.
Publicado: (2024)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
por: Qiu, Jielin, et al.
Publicado: (2025)
por: Qiu, Jielin, et al.
Publicado: (2025)
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook
por: Yu, Xinlei, et al.
Publicado: (2026)
por: Yu, Xinlei, et al.
Publicado: (2026)
MatSciBench: Benchmarking the Reasoning Ability of Large Language Models in Materials Science
por: Zhang, Junkai, et al.
Publicado: (2025)
por: Zhang, Junkai, et al.
Publicado: (2025)
ProBench: Benchmarking GUI Agents with Accurate Process Information
por: Yang, Leyang, et al.
Publicado: (2025)
por: Yang, Leyang, et al.
Publicado: (2025)
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
por: Liu, Junqi, et al.
Publicado: (2025)
por: Liu, Junqi, et al.
Publicado: (2025)
Ejemplares similares
-
Tailoring the Curriculum: Student-Centered Reasoning Distillation via Dynamic Data-Model Compatibility
por: Huang, Jiahao, et al.
Publicado: (2026) -
JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models
por: Jiang, Junfeng, et al.
Publicado: (2024) -
InnovatorBench: Evaluating Agents' Ability to Conduct Innovative LLM Research
por: Wu, Yunze, et al.
Publicado: (2025) -
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents
por: Costarelli, Anthony, et al.
Publicado: (2024) -
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
por: Wang, Ruipeng, et al.
Publicado: (2026)