GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Diao, Lingxiao, Xu, Xinyue, Sun, Wanxuan, Yang, Cheng, Zhang, Zhuosheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
por: Ghazarian, Sarik, et al.
Publicado: (2025)
por: Ghazarian, Sarik, et al.
Publicado: (2025)
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
por: Xiao, Ruixuan, et al.
Publicado: (2024)
por: Xiao, Ruixuan, et al.
Publicado: (2024)
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
por: Guo, Yuan, et al.
Publicado: (2025)
por: Guo, Yuan, et al.
Publicado: (2025)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024)
por: Deng, Shihan, et al.
Publicado: (2024)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
por: Li, Haitao, et al.
Publicado: (2024)
por: Li, Haitao, et al.
Publicado: (2024)
ChemSafetyBench: Benchmarking LLM Safety on Chemistry Domain
por: Zhao, Haochen, et al.
Publicado: (2024)
por: Zhao, Haochen, et al.
Publicado: (2024)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
por: Zhang, Shaoqing, et al.
Publicado: (2024)
por: Zhang, Shaoqing, et al.
Publicado: (2024)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
por: Cheng, Xiang, et al.
Publicado: (2026)
por: Cheng, Xiang, et al.
Publicado: (2026)
Smoothing Grounding and Reasoning for MLLM-Powered GUI Agents with Query-Oriented Pivot Tasks
por: Wu, Zongru, et al.
Publicado: (2025)
por: Wu, Zongru, et al.
Publicado: (2025)
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
por: Wang, Jiangyuan, et al.
Publicado: (2025)
por: Wang, Jiangyuan, et al.
Publicado: (2025)
GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations
por: Yang, Jingbo, et al.
Publicado: (2026)
por: Yang, Jingbo, et al.
Publicado: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
por: Yuan, Tongxin, et al.
Publicado: (2024)
por: Yuan, Tongxin, et al.
Publicado: (2024)
CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
por: Ma, Xinbei, et al.
Publicado: (2024)
por: Ma, Xinbei, et al.
Publicado: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
por: Zhang, Zhuosheng, et al.
Publicado: (2023)
por: Zhang, Zhuosheng, et al.
Publicado: (2023)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
por: Zhang, Zhexin, et al.
Publicado: (2024)
por: Zhang, Zhexin, et al.
Publicado: (2024)
ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
por: Song, Yuanyi, et al.
Publicado: (2025)
por: Song, Yuanyi, et al.
Publicado: (2025)
DataSciBench: An LLM Agent Benchmark for Data Science
por: Zhang, Dan, et al.
Publicado: (2025)
por: Zhang, Dan, et al.
Publicado: (2025)
MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents
por: Du, Yiming, et al.
Publicado: (2025)
por: Du, Yiming, et al.
Publicado: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
por: Perlitz, Yotam, et al.
Publicado: (2024)
por: Perlitz, Yotam, et al.
Publicado: (2024)
CryptoBench: A Dynamic Benchmark for Expert-Level Evaluation of LLM Agents in Cryptocurrency
por: Guo, Jiacheng, et al.
Publicado: (2025)
por: Guo, Jiacheng, et al.
Publicado: (2025)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
por: Wen, Bosi, et al.
Publicado: (2026)
por: Wen, Bosi, et al.
Publicado: (2026)
FHIR-AgentBench: Benchmarking LLM Agents for Realistic Interoperable EHR Question Answering
por: Lee, Gyubok, et al.
Publicado: (2025)
por: Lee, Gyubok, et al.
Publicado: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
por: He, Wei, et al.
Publicado: (2025)
por: He, Wei, et al.
Publicado: (2025)
Caution for the Environment: Multimodal LLM Agents are Susceptible to Environmental Distractions
por: Ma, Xinbei, et al.
Publicado: (2024)
por: Ma, Xinbei, et al.
Publicado: (2024)
OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
por: Ding, Deming, et al.
Publicado: (2026)
por: Ding, Deming, et al.
Publicado: (2026)
DOCBENCH: A Benchmark for Evaluating LLM-based Document Reading Systems
por: Zou, Anni, et al.
Publicado: (2024)
por: Zou, Anni, et al.
Publicado: (2024)
PodBench: A Comprehensive Benchmark for Instruction-Aware Audio-Oriented Podcast Script Generation
por: Xu, Chenning, et al.
Publicado: (2026)
por: Xu, Chenning, et al.
Publicado: (2026)
StreamBench: Towards Benchmarking Continuous Improvement of Language Agents
por: Wu, Cheng-Kuang, et al.
Publicado: (2024)
por: Wu, Cheng-Kuang, et al.
Publicado: (2024)
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
por: Nguyen, Bang, et al.
Publicado: (2026)
por: Nguyen, Bang, et al.
Publicado: (2026)
LCTG Bench: LLM Controlled Text Generation Benchmark
por: Kurihara, Kentaro, et al.
Publicado: (2025)
por: Kurihara, Kentaro, et al.
Publicado: (2025)
Flooding Spread of Manipulated Knowledge in LLM-Based Multi-Agent Communities
por: Ju, Tianjie, et al.
Publicado: (2024)
por: Ju, Tianjie, et al.
Publicado: (2024)
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
por: Wang, Zhenting, et al.
Publicado: (2025)
por: Wang, Zhenting, et al.
Publicado: (2025)
CO-Bench: Benchmarking Language Model Agents in Algorithm Search for Combinatorial Optimization
por: Sun, Weiwei, et al.
Publicado: (2025)
por: Sun, Weiwei, et al.
Publicado: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
por: Tu, Xinming, et al.
Publicado: (2026)
por: Tu, Xinming, et al.
Publicado: (2026)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
por: Afzal, Anum, et al.
Publicado: (2025)
por: Afzal, Anum, et al.
Publicado: (2025)
EduBench: A Comprehensive Benchmarking Dataset for Evaluating Large Language Models in Diverse Educational Scenarios
por: Xu, Bin, et al.
Publicado: (2025)
por: Xu, Bin, et al.
Publicado: (2025)
ISD-Agent-Bench: A Comprehensive Benchmark for Evaluating LLM-based Instructional Design Agents
por: Jeon, YoungHoon, et al.
Publicado: (2026)
por: Jeon, YoungHoon, et al.
Publicado: (2026)
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
por: Jiang, Yuxin, et al.
Publicado: (2023)
por: Jiang, Yuxin, et al.
Publicado: (2023)
GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
por: Wu, Zheng, et al.
Publicado: (2025)
por: Wu, Zheng, et al.
Publicado: (2025)
ENPMR-Bench: Benchmarking Proactive Memory Retrieval for Emotional Support Agents
por: Fu, Xing, et al.
Publicado: (2026)
por: Fu, Xing, et al.
Publicado: (2026)
Ejemplares similares
-
TOD-ProcBench: Benchmarking Complex Instruction-Following in Task-Oriented Dialogues
por: Ghazarian, Sarik, et al.
Publicado: (2025) -
FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
por: Xiao, Ruixuan, et al.
Publicado: (2024) -
Atomic-to-Compositional Generalization for Mobile Agents with A New Benchmark and Scheduling System
por: Guo, Yuan, et al.
Publicado: (2025) -
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
por: Deng, Shihan, et al.
Publicado: (2024) -
LegalAgentBench: Evaluating LLM Agents in Legal Domain
por: Li, Haitao, et al.
Publicado: (2024)