FlowBench: Revisiting and Benchmarking Workflow-Guided Planning for LLM-based Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Ruixuan, Ma, Wentao, Wang, Ke, Wu, Yuchuan, Zhao, Junbo, Wang, Haobo, Huang, Fei, Li, Yongbin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlowBench: A Large Scale Benchmark for Flow Simulation over Complex Geometries
by: Tali, Ronak, et al.
Published: (2024)
by: Tali, Ronak, et al.
Published: (2024)
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025)
by: Tian, Jiaming, et al.
Published: (2025)
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025)
by: Liao, Chonghua, et al.
Published: (2025)
SDPO: Segment-Level Direct Preference Optimization for Social Agents
by: Kong, Aobo, et al.
Published: (2025)
by: Kong, Aobo, et al.
Published: (2025)
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)
by: Huang, Ruixuan, et al.
Published: (2025)
SpokenWOZ: A Large-Scale Speech-Text Benchmark for Spoken Task-Oriented Dialogue Agents
by: Si, Shuzheng, et al.
Published: (2023)
by: Si, Shuzheng, et al.
Published: (2023)
DataGovBench: Benchmarking LLM Agents for Real-World Data Governance Workflows
by: Liu, Zhou, et al.
Published: (2025)
by: Liu, Zhou, et al.
Published: (2025)
EPO: Explicit Policy Optimization for Strategic Reasoning in LLMs via Reinforcement Learning
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
Aligning Logits Generatively for Principled Black-Box Knowledge Distillation
by: Ma, Jing, et al.
Published: (2022)
by: Ma, Jing, et al.
Published: (2022)
RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
by: Wu, Pengzuo, et al.
Published: (2025)
by: Wu, Pengzuo, et al.
Published: (2025)
On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey
by: Long, Lin, et al.
Published: (2024)
by: Long, Lin, et al.
Published: (2024)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
Agentic Reinforcement Learning with Implicit Step Rewards
by: Liu, Xiaoqian, et al.
Published: (2025)
by: Liu, Xiaoqian, et al.
Published: (2025)
RECOST: External Knowledge Guided Data-efficient Instruction Tuning
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
by: Fu, Yuchuan, et al.
Published: (2025)
by: Fu, Yuchuan, et al.
Published: (2025)
SPA++: Generalized Graph Spectral Alignment for Versatile Domain Adaptation
by: Xiao, Zhiqing, et al.
Published: (2025)
by: Xiao, Zhiqing, et al.
Published: (2025)
Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
by: Zhang, Hanrong, et al.
Published: (2024)
by: Zhang, Hanrong, et al.
Published: (2024)
CPO: Addressing Reward Ambiguity in Role-playing Dialogue via Comparative Policy Optimization
by: Ye, Xinge, et al.
Published: (2025)
by: Ye, Xinge, et al.
Published: (2025)
ShoppingBench: A Real-World Intent-Grounded Shopping Benchmark for LLM-based Agents
by: Wang, Jiangyuan, et al.
Published: (2025)
by: Wang, Jiangyuan, et al.
Published: (2025)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
AgentRecBench: Benchmarking LLM Agent-based Personalized Recommender Systems
by: Shang, Yu, et al.
Published: (2025)
by: Shang, Yu, et al.
Published: (2025)
ScoreFlow: Mastering LLM Agent Workflows via Score-based Preference Optimization
by: Wang, Yinjie, et al.
Published: (2025)
by: Wang, Yinjie, et al.
Published: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
Adaptive Social Learning via Mode Policy Optimization for Language Agents
by: Wang, Minzheng, et al.
Published: (2025)
by: Wang, Minzheng, et al.
Published: (2025)
Mobile-Bench: An Evaluation Benchmark for LLM-based Mobile Agents
by: Deng, Shihan, et al.
Published: (2024)
by: Deng, Shihan, et al.
Published: (2024)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
CodeFlowBench: A Multi-turn, Iterative Benchmark for Complex Code Generation
by: Wang, Sizhe, et al.
Published: (2025)
by: Wang, Sizhe, et al.
Published: (2025)
GuideBench: Benchmarking Domain-Oriented Guideline Following for LLM Agents
by: Diao, Lingxiao, et al.
Published: (2025)
by: Diao, Lingxiao, et al.
Published: (2025)
MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows
by: Shen, Weixiang, et al.
Published: (2026)
by: Shen, Weixiang, et al.
Published: (2026)
Self-Organizing Agent Network for LLM-based Workflow Automation
by: Xiong, Yiming, et al.
Published: (2025)
by: Xiong, Yiming, et al.
Published: (2025)
Data Contamination Calibration for Black-box LLMs
by: Ye, Wentao, et al.
Published: (2024)
by: Ye, Wentao, et al.
Published: (2024)
BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
Navigate Complex Physical Worlds via Geometrically Constrained LLM
by: Huang, Yongqiang, et al.
Published: (2024)
by: Huang, Yongqiang, et al.
Published: (2024)
Energy-based Automated Model Evaluation
by: Peng, Ru, et al.
Published: (2024)
by: Peng, Ru, et al.
Published: (2024)
EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents
by: Dong, Xuan, et al.
Published: (2026)
by: Dong, Xuan, et al.
Published: (2026)
Benchmarking LLM Agents for Wealth-Management Workflows
by: Milsom, Rory
Published: (2025)
by: Milsom, Rory
Published: (2025)
Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation
by: Xia, Mingxuan, et al.
Published: (2025)
by: Xia, Mingxuan, et al.
Published: (2025)
GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning
by: Cheng, Xiang, et al.
Published: (2026)
by: Cheng, Xiang, et al.
Published: (2026)
OdysseyBench: Evaluating LLM Agents on Long-Horizon Complex Office Application Workflows
by: Wang, Weixuan, et al.
Published: (2025)
by: Wang, Weixuan, et al.
Published: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
Similar Items
-
FlowBench: A Large Scale Benchmark for Flow Simulation over Complex Geometries
by: Tali, Ronak, et al.
Published: (2024) -
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
by: Tian, Jiaming, et al.
Published: (2025) -
MOA: Multi-Objective Alignment for Role-Playing Agents
by: Liao, Chonghua, et al.
Published: (2025) -
SDPO: Segment-Level Direct Preference Optimization for Social Agents
by: Kong, Aobo, et al.
Published: (2025) -
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
by: Huang, Ruixuan, et al.
Published: (2025)