AWARE-US: Preference-Aware Infeasibility Resolution in Tool-Calling Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Kurmaz, Mehmet
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910036393984000
author Kurmaz, Mehmet
author_facet Kurmaz, Mehmet
contents Tool-calling conversational agents querying structured databases often face two linked failures: underspecification (missing constraints needed for a precise query) andinfeasibility (a fully specified query returns anemptyset). Prior systems often respond with "no results" or apply ad hoc relaxations, which can violate user intent by discarding highly valued requirements. Wecast infeasibility handling as preference-aware query repair: when a query is unsatisfiable, the agent should relax the least important constraints. We propose three LLM-based methods to infer relative constraint importance from dialogue: (1) local weighting, (2) global one-shot weighting, and (3) pairwise ranking. Across extensive experiments in car recommendation, the local-weighting method trained with supervised fine-tuning and direct preference optimization best aligns with user preferences (48%), while global weighting achieves the highest correct-relaxation accuracy (56%); all three outperform prior infeasibility-resolution basel. We also introduce AWARE-US, a benchmark of 120+ persona-grounded queries requiring agents to (i) disambiguate a base request via conversa tion and (ii) resolve infeasibility in a way consistent with persona-implied preferences. For code refer to Github: https://github.com/mhtkrmz/Infeasible-task and the dataset is available on Hugging Face
format Preprint
id arxiv_https___arxiv_org_abs_2601_02643
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AWARE-US: Preference-Aware Infeasibility Resolution in Tool-Calling Agents
Kurmaz, Mehmet
Artificial Intelligence
Tool-calling conversational agents querying structured databases often face two linked failures: underspecification (missing constraints needed for a precise query) andinfeasibility (a fully specified query returns anemptyset). Prior systems often respond with "no results" or apply ad hoc relaxations, which can violate user intent by discarding highly valued requirements. Wecast infeasibility handling as preference-aware query repair: when a query is unsatisfiable, the agent should relax the least important constraints. We propose three LLM-based methods to infer relative constraint importance from dialogue: (1) local weighting, (2) global one-shot weighting, and (3) pairwise ranking. Across extensive experiments in car recommendation, the local-weighting method trained with supervised fine-tuning and direct preference optimization best aligns with user preferences (48%), while global weighting achieves the highest correct-relaxation accuracy (56%); all three outperform prior infeasibility-resolution basel. We also introduce AWARE-US, a benchmark of 120+ persona-grounded queries requiring agents to (i) disambiguate a base request via conversa tion and (ii) resolve infeasibility in a way consistent with persona-implied preferences. For code refer to Github: https://github.com/mhtkrmz/Infeasible-task and the dataset is available on Hugging Face
title AWARE-US: Preference-Aware Infeasibility Resolution in Tool-Calling Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2601.02643