SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zekun, Huang, Shinda, Wang, Jiangtian, Zhang, Nathan, Antoniades, Antonis, Hua, Wenyue, Zhu, Kaijie, Zeng, Sirui, Wang, Chi, Wang, William Yang, Yan, Xifeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
by: Amayuelas, Alfonso, et al.
Published: (2024)
by: Amayuelas, Alfonso, et al.
Published: (2024)
ADL: A Declarative Language for Agent-Based Chatbots
by: Zeng, Sirui, et al.
Published: (2025)
by: Zeng, Sirui, et al.
Published: (2025)
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
by: Antoniades, Antonis, et al.
Published: (2024)
by: Antoniades, Antonis, et al.
Published: (2024)
Self-Resource Allocation in Multi-Agent LLM Systems
by: Amayuelas, Alfonso, et al.
Published: (2025)
by: Amayuelas, Alfonso, et al.
Published: (2025)
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
by: Weng, Zhaotian, et al.
Published: (2026)
by: Weng, Zhaotian, et al.
Published: (2026)
Neuroformer: Multimodal and Multitask Generative Pretraining for Brain Data
by: Antoniades, Antonis, et al.
Published: (2023)
by: Antoniades, Antonis, et al.
Published: (2023)
MAGPIE: A dataset for Multi-AGent contextual PrIvacy Evaluation
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
Dynamic Speculative Agent Planning
by: Guan, Yilin, et al.
Published: (2025)
by: Guan, Yilin, et al.
Published: (2025)
Generalization v.s. Memorization: Tracing Language Models' Capabilities Back to Pretraining Data
by: Wang, Xinyi, et al.
Published: (2024)
by: Wang, Xinyi, et al.
Published: (2024)
Language Control Diffusion: Efficiently Scaling through Space, Time, and Tasks
by: Zhang, Edwin, et al.
Published: (2022)
by: Zhang, Edwin, et al.
Published: (2022)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Agent-S: LLM Agentic workflow to automate Standard Operating Procedures
by: Kulkarni, Mandar
Published: (2025)
by: Kulkarni, Mandar
Published: (2025)
Formal-LLM: Integrating Formal Language and Natural Language for Controllable LLM-based Agents
by: Li, Zelong, et al.
Published: (2024)
by: Li, Zelong, et al.
Published: (2024)
Semantic minimalism and the continuous nature of polysemy
by: Jiangtian Li
Published: (2024)
by: Jiangtian Li
Published: (2024)
SOP-Maze: Evaluating Large Language Models on Complicated Business Standard Operating Procedures
by: Wang, Jiaming, et al.
Published: (2025)
by: Wang, Jiaming, et al.
Published: (2025)
Dynamic Evaluation of Large Language Models by Meta Probing Agents
by: Zhu, Kaijie, et al.
Published: (2024)
by: Zhu, Kaijie, et al.
Published: (2024)
Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Multiple kutane epitheloide angiomatöse Knötchen
by: Sirui Hua, et al.
Published: (2025)
by: Sirui Hua, et al.
Published: (2025)
High Dimensional Procedural Content Generation
by: Xu, Kaijie, et al.
Published: (2026)
by: Xu, Kaijie, et al.
Published: (2026)
MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
by: Zhu, Kaijie, et al.
Published: (2025)
by: Zhu, Kaijie, et al.
Published: (2025)
Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
by: Hua, Wenyue, et al.
Published: (2026)
by: Hua, Wenyue, et al.
Published: (2026)
AIOS: LLM Agent Operating System
by: Mei, Kai, et al.
Published: (2024)
by: Mei, Kai, et al.
Published: (2024)
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models
by: Pu, Xiao, et al.
Published: (2025)
by: Pu, Xiao, et al.
Published: (2025)
MAGPIE: A benchmark for Multi-AGent contextual PrIvacy Evaluation
by: Juneja, Gurusha, et al.
Published: (2025)
by: Juneja, Gurusha, et al.
Published: (2025)
SOPRAG: Multi-view Graph Experts Retrieval for Industrial Standard Operating Procedures
by: Lin, Liangtao, et al.
Published: (2026)
by: Lin, Liangtao, et al.
Published: (2026)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Direct Multi-Token Decoding
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement Learning
by: Wang, Danqing, et al.
Published: (2024)
by: Wang, Danqing, et al.
Published: (2024)
What if LLMs Have Different World Views: Simulating Alien Civilizations with LLM-based Agents
by: Xue, Zhaoqian, et al.
Published: (2024)
by: Xue, Zhaoqian, et al.
Published: (2024)
OpenP5: An Open-Source Platform for Developing, Training, and Evaluating LLM-based Recommender Systems
by: Xu, Shuyuan, et al.
Published: (2023)
by: Xu, Shuyuan, et al.
Published: (2023)
AgentReview: Exploring Peer Review Dynamics with LLM Agents
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
Bot or Human? Detecting ChatGPT Imposters with A Single Question
by: Wang, Hong, et al.
Published: (2023)
by: Wang, Hong, et al.
Published: (2023)
Learning to Lie: Reinforcement Learning Attacks Damage Human-AI Teams and Teams of LLMs
by: Musaffar, Abed Kareem, et al.
Published: (2025)
by: Musaffar, Abed Kareem, et al.
Published: (2025)
Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Can Online GenAI Discussion Serve as Bellwether for Labor Market Shifts?
by: Cao, Shurui, et al.
Published: (2025)
by: Cao, Shurui, et al.
Published: (2025)
Multi‐Agent Learning With Hierarchical Biomechanical Priors for Efficient 3D Human Pose Estimation in Virtual Reality
by: Xingquan Cai, et al.
Published: (2025)
by: Xingquan Cai, et al.
Published: (2025)
Design Study for Project on Standard Operating Procedures for Technical Library Services.
by: Libbey, Miles A., et al.
Published: (1970)
by: Libbey, Miles A., et al.
Published: (1970)
PromptBench: A Unified Library for Evaluation of Large Language Models
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
Metal‐Organic Nanosheet Gels: Hierarchically Porous Materials for Selective Loading and Differential Release
by: Jiangtian Tan, et al.
Published: (2025)
by: Jiangtian Tan, et al.
Published: (2025)
Probing the Representational Structure of Regular Polysemy via Sense Analogy Questions: Insights from Contextual Word Vectors
by: Jiangtian Li, et al.
Published: (2024)
by: Jiangtian Li, et al.
Published: (2024)
Similar Items
-
MultiAgent Collaboration Attack: Investigating Adversarial Attacks in Large Language Model Collaborations via Debate
by: Amayuelas, Alfonso, et al.
Published: (2024) -
ADL: A Declarative Language for Agent-Based Chatbots
by: Zeng, Sirui, et al.
Published: (2025) -
SWE-Search: Enhancing Software Agents with Monte Carlo Tree Search and Iterative Refinement
by: Antoniades, Antonis, et al.
Published: (2024) -
Self-Resource Allocation in Multi-Agent LLM Systems
by: Amayuelas, Alfonso, et al.
Published: (2025) -
Group-Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
by: Weng, Zhaotian, et al.
Published: (2026)