LHAW: Controllable Underspecification for Long-Horizon Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pu, George, Lee, Michael S., Sehwag, Udari Madhushani, Lee, David J., Zhu, Bryan, Maurya, Yash, Raghavendra, Mohit, Xue, Yuan, Denton, Samuel Marc
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917354523328512
author Pu, George
Lee, Michael S.
Sehwag, Udari Madhushani
Lee, David J.
Zhu, Bryan
Maurya, Yash
Raghavendra, Mohit
Xue, Yuan
Denton, Samuel Marc
author_facet Pu, George
Lee, Michael S.
Sehwag, Udari Madhushani
Lee, David J.
Zhu, Bryan
Maurya, Yash
Raghavendra, Mohit
Xue, Yuan
Denton, Samuel Marc
contents Long-horizon workflow agents that operate effectively over extended periods are essential for truly autonomous systems. Their reliable execution critically depends on the ability to reason through ambiguous situations in which clarification seeking is necessary to ensure correct task execution. However, progress is limited by the lack of scalable, task-agnostic frameworks for systematically curating and measuring the impact of ambiguity across custom workflows. We address this gap by introducing LHAW (Long-Horizon Augmented Workflows), a modular, dataset-agnostic synthetic pipeline that transforms any well-specified task into controllable underspecified variants by systematically removing information across four dimensions - Goals, Constraints, Inputs, and Context - at configurable severity levels. Unlike approaches that rely on LLM predictions of ambiguity, LHAW validates variants through empirical agent trials, classifying them as outcome-critical, divergent, or benign based on observed terminal state divergence. We release 285 task variants from TheAgentCompany, SWE-Bench Pro and MCP-Atlas according to our taxonomy alongside formal analysis measuring how current agents detect, reason about, and resolve underspecification across ambiguous settings. LHAW provides the first systematic framework for cost-sensitive evaluation of agent clarification behavior in long-horizon settings, enabling development of reliable autonomous systems.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10525
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LHAW: Controllable Underspecification for Long-Horizon Tasks
Pu, George
Lee, Michael S.
Sehwag, Udari Madhushani
Lee, David J.
Zhu, Bryan
Maurya, Yash
Raghavendra, Mohit
Xue, Yuan
Denton, Samuel Marc
Computation and Language
Artificial Intelligence
Machine Learning
Long-horizon workflow agents that operate effectively over extended periods are essential for truly autonomous systems. Their reliable execution critically depends on the ability to reason through ambiguous situations in which clarification seeking is necessary to ensure correct task execution. However, progress is limited by the lack of scalable, task-agnostic frameworks for systematically curating and measuring the impact of ambiguity across custom workflows. We address this gap by introducing LHAW (Long-Horizon Augmented Workflows), a modular, dataset-agnostic synthetic pipeline that transforms any well-specified task into controllable underspecified variants by systematically removing information across four dimensions - Goals, Constraints, Inputs, and Context - at configurable severity levels. Unlike approaches that rely on LLM predictions of ambiguity, LHAW validates variants through empirical agent trials, classifying them as outcome-critical, divergent, or benign based on observed terminal state divergence. We release 285 task variants from TheAgentCompany, SWE-Bench Pro and MCP-Atlas according to our taxonomy alongside formal analysis measuring how current agents detect, reason about, and resolve underspecification across ambiguous settings. LHAW provides the first systematic framework for cost-sensitive evaluation of agent clarification behavior in long-horizon settings, enabling development of reliable autonomous systems.
title LHAW: Controllable Underspecification for Long-Horizon Tasks
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2602.10525