Saved in:
| Main Authors: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2602.07549 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866914313996861440 |
|---|---|
| author | Ko, Dayoon Kim, Jihyuk Kim, Sohyeon Park, Haeju Lee, Dahyun Kim, Gunhee Lee, Moontae Lee, Kyungjae |
| author_facet | Ko, Dayoon Kim, Jihyuk Kim, Sohyeon Park, Haeju Lee, Dahyun Kim, Gunhee Lee, Moontae Lee, Kyungjae |
| contents | Recent search agents leverage multi-turn reasoning and search tools to achieve strong performance on multi-hop and long-horizon benchmarks. Yet it remains unclear whether they reliably reason across all requirements by tracking, verifying, and maintaining multiple conditions in these questions. We study this capability under multi-constraint problems, where valid answers must satisfy several constraints simultaneously. We find that illusory completion frequently occurs, wherein agents believe tasks are complete despite unresolved or violated constraints, leading to underverified answers. To diagnose this behavior, we introduce the Epistemic Ledger, an evaluation framework that tracks evidential support and agents' beliefs for each constraint throughout multi-turn reasoning. Our analysis reveals four recurring failure patterns: bare assertions, overlooked refutations, stagnation, and premature exit. Motivated by these findings, we examine whether explicit constraint-state tracking during execution mitigates these failures via LiveLedger, an inference-time tracker. This simple intervention consistently improves performance, substantially reducing underverified answers (by up to 26.5%) and improving overall accuracy (by up to 11.6%) on multi-constraint problems. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_07549 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | When Is Enough Not Enough? Illusory Completion in Search Agents Ko, Dayoon Kim, Jihyuk Kim, Sohyeon Park, Haeju Lee, Dahyun Kim, Gunhee Lee, Moontae Lee, Kyungjae Artificial Intelligence Computation and Language Recent search agents leverage multi-turn reasoning and search tools to achieve strong performance on multi-hop and long-horizon benchmarks. Yet it remains unclear whether they reliably reason across all requirements by tracking, verifying, and maintaining multiple conditions in these questions. We study this capability under multi-constraint problems, where valid answers must satisfy several constraints simultaneously. We find that illusory completion frequently occurs, wherein agents believe tasks are complete despite unresolved or violated constraints, leading to underverified answers. To diagnose this behavior, we introduce the Epistemic Ledger, an evaluation framework that tracks evidential support and agents' beliefs for each constraint throughout multi-turn reasoning. Our analysis reveals four recurring failure patterns: bare assertions, overlooked refutations, stagnation, and premature exit. Motivated by these findings, we examine whether explicit constraint-state tracking during execution mitigates these failures via LiveLedger, an inference-time tracker. This simple intervention consistently improves performance, substantially reducing underverified answers (by up to 26.5%) and improving overall accuracy (by up to 11.6%) on multi-constraint problems. |
| title | When Is Enough Not Enough? Illusory Completion in Search Agents |
| topic | Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2602.07549 |