When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Henry Peng, Miao, Chunyu, Huang, Wei-Chieh, Chen, Yankai, Zhou, Yue, Zhang, Hanrong, Wu, Yaozu, Fang, Liancheng, Gu, Zhengyao, Zhang, Zhen, Zheng, Kening, Wang, Fangxin, Nian, Yi, Li, Shanghao, Fan, Wenzhe, He, Langzhou, Zhang, Weizhi, Liu, Xue, Yu, Philip S. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
by: Huang, Wei-Chieh, et al.
Published: (2025)
by: Huang, Wei-Chieh, et al.
Published: (2025)
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
by: Gu, Zhengyao, et al.
Published: (2025)
by: Gu, Zhengyao, et al.
Published: (2025)
TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency
by: Zou, Henry Peng, et al.
Published: (2025)
by: Zou, Henry Peng, et al.
Published: (2025)
Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
by: Chen, Yankai, et al.
Published: (2025)
by: Chen, Yankai, et al.
Published: (2025)
TabGen-ICL: Residual-Aware In-Context Example Selection for Tabular Data Generation
by: Fang, Liancheng, et al.
Published: (2025)
by: Fang, Liancheng, et al.
Published: (2025)
Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models
by: Fang, Liancheng, et al.
Published: (2026)
by: Fang, Liancheng, et al.
Published: (2026)
Multi-Agent Autonomous Driving Systems with Large Language Models: A Survey of Recent Advances
by: Wu, Yaozu, et al.
Published: (2025)
by: Wu, Yaozu, et al.
Published: (2025)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
by: Wang, Fangxin, et al.
Published: (2026)
by: Wang, Fangxin, et al.
Published: (2026)
TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System
by: Fan, Wenzhe, et al.
Published: (2026)
by: Fan, Wenzhe, et al.
Published: (2026)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
by: Wu, Yaozu, et al.
Published: (2025)
by: Wu, Yaozu, et al.
Published: (2025)
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
by: Zhang, Hanrong, et al.
Published: (2026)
by: Zhang, Hanrong, et al.
Published: (2026)
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
by: Miao, Chunyu, et al.
Published: (2025)
by: Miao, Chunyu, et al.
Published: (2025)
MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample Selection
by: Fang, Liancheng, et al.
Published: (2025)
by: Fang, Liancheng, et al.
Published: (2025)
Do We Really Need Graph Convolution During Training? Light Post-Training Graph-ODE for Efficient Recommendation
by: Zhang, Weizhi, et al.
Published: (2024)
by: Zhang, Weizhi, et al.
Published: (2024)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
by: He, Langzhou, et al.
Published: (2026)
by: He, Langzhou, et al.
Published: (2026)
GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
by: Wu, Zhaofen, et al.
Published: (2026)
by: Wu, Zhaofen, et al.
Published: (2026)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
by: Zheng, Kening, et al.
Published: (2026)
by: Zheng, Kening, et al.
Published: (2026)
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
by: Zhang, Weizhi, et al.
Published: (2025)
by: Zhang, Weizhi, et al.
Published: (2025)
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
by: Chen, Yankai, et al.
Published: (2026)
by: Chen, Yankai, et al.
Published: (2026)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
by: Zou, Henry Peng, et al.
Published: (2024)
by: Zou, Henry Peng, et al.
Published: (2024)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
by: Guo, Zhen, et al.
Published: (2025)
by: Guo, Zhen, et al.
Published: (2025)
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
by: Li, Chengze, et al.
Published: (2026)
by: Li, Chengze, et al.
Published: (2026)
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
by: Gu, Zhengyao, et al.
Published: (2026)
by: Gu, Zhengyao, et al.
Published: (2026)
From Uncertainty to Clarity: Uncertainty-Guided Class-Incremental Learning for Limited Biomedical Samples via Semantic Expansion
by: Yao, Yifei, et al.
Published: (2024)
by: Yao, Yifei, et al.
Published: (2024)
Enabling Deterministic User-Level Interrupts in Real-Time Processors via Hardware Extension
by: Yang, Hongbin, et al.
Published: (2026)
by: Yang, Hongbin, et al.
Published: (2026)
MANSY: Generalizing Neural Adaptive Immersive Video Streaming With Ensemble and Representation Learning
by: Wu, Duo, et al.
Published: (2023)
by: Wu, Duo, et al.
Published: (2023)
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
by: Zhang, Weizhi, et al.
Published: (2026)
by: Zhang, Weizhi, et al.
Published: (2026)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
by: Li, Shanghao, et al.
Published: (2026)
by: Li, Shanghao, et al.
Published: (2026)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
by: Li, Shanghao, et al.
Published: (2025)
by: Li, Shanghao, et al.
Published: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
by: Zhao, Bingchen, et al.
Published: (2026)
by: Zhao, Bingchen, et al.
Published: (2026)
Handling The Non-Smooth Challenge in Tensor SVD: A Multi-Objective Tensor Recovery Framework
by: Zheng, Jingjing, et al.
Published: (2023)
by: Zheng, Jingjing, et al.
Published: (2023)
Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
by: He, Langzhou, et al.
Published: (2025)
by: He, Langzhou, et al.
Published: (2025)
LawMind: A Law-Driven Paradigm for Discovering Analytical Solutions to Partial Differential Equations
by: Zheng, Min-Yi, et al.
Published: (2026)
by: Zheng, Min-Yi, et al.
Published: (2026)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
by: Yao, Yifei, et al.
Published: (2025)
by: Yao, Yifei, et al.
Published: (2025)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models
by: Zhang, Liancheng, et al.
Published: (2026)
by: Zhang, Liancheng, et al.
Published: (2026)
Tough and Recyclable Poly(phthalazinone ether sulfone ketone) Aerogel Fibers
by: Chen Zhuo, et al.
Published: (2026)
by: Chen Zhuo, et al.
Published: (2026)
Similar Items
-
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
by: Zou, Henry Peng, et al.
Published: (2025) -
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
by: Huang, Wei-Chieh, et al.
Published: (2025) -
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
by: Zou, Henry Peng, et al.
Published: (2025) -
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
by: Gu, Zhengyao, et al.
Published: (2025) -
TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency
by: Zou, Henry Peng, et al.
Published: (2025)