SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiaming, Tang, Zhe, Jin, Zehao, Chen, Hefei, Jin, Yilin, Ding, Peng, Li, Xiaoyu, Cao, Xuezhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912822273769472
author Wang, Jiaming
Tang, Zhe
Jin, Zehao
Chen, Hefei
Jin, Yilin
Ding, Peng
Li, Xiaoyu
Cao, Xuezhi
author_facet Wang, Jiaming
Tang, Zhe
Jin, Zehao
Chen, Hefei
Jin, Yilin
Ding, Peng
Li, Xiaoyu
Cao, Xuezhi
contents As large language models (LLMs) are widely deployed as domain-specific agents, many benchmarks have been proposed to evaluate their ability to follow instructions and make decisions in real-world scenarios. However, business scenarios often involve complex standard operating procedures (SOPs), and the evaluation of LLM capabilities in such contexts has not been fully explored. To bridge this gap, we propose SOP-Maze, a benchmark constructed from real-world business data and adapted into a collection of 397 instances and 3422 subtasks from 23 complex SOP scenarios. We further categorize SOP tasks into two broad classes: Lateral Root System (LRS), representing wide-option tasks that demand precise selection; and Heart Root System (HRS), which emphasizes deep logical reasoning with complex branches. Extensive experiments reveal that nearly all state-of-the-art models struggle with SOP-Maze. We conduct a comprehensive analysis and identify three key error categories: (i) route blindness: difficulty following procedures; (ii) conversational fragility: inability to handle real dialogue nuances; and (iii) calculation errors: mistakes in time or arithmetic reasoning under complex contexts. The systematic study explores LLM performance across SOP tasks that challenge both breadth and depth, offering new insights for improving model capabilities. We have open-sourced our work on: https://github.com/meituan-longcat/SOP-Maze.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08942
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
Wang, Jiaming
Tang, Zhe
Jin, Zehao
Chen, Hefei
Jin, Yilin
Ding, Peng
Li, Xiaoyu
Cao, Xuezhi
Computation and Language
As large language models (LLMs) are widely deployed as domain-specific agents, many benchmarks have been proposed to evaluate their ability to follow instructions and make decisions in real-world scenarios. However, business scenarios often involve complex standard operating procedures (SOPs), and the evaluation of LLM capabilities in such contexts has not been fully explored. To bridge this gap, we propose SOP-Maze, a benchmark constructed from real-world business data and adapted into a collection of 397 instances and 3422 subtasks from 23 complex SOP scenarios. We further categorize SOP tasks into two broad classes: Lateral Root System (LRS), representing wide-option tasks that demand precise selection; and Heart Root System (HRS), which emphasizes deep logical reasoning with complex branches. Extensive experiments reveal that nearly all state-of-the-art models struggle with SOP-Maze. We conduct a comprehensive analysis and identify three key error categories: (i) route blindness: difficulty following procedures; (ii) conversational fragility: inability to handle real dialogue nuances; and (iii) calculation errors: mistakes in time or arithmetic reasoning under complex contexts. The systematic study explores LLM performance across SOP tasks that challenge both breadth and depth, offering new insights for improving model capabilities. We have open-sourced our work on: https://github.com/meituan-longcat/SOP-Maze.
title SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
topic Computation and Language
url https://arxiv.org/abs/2510.08942