When Users Change Their Mind: Evaluating Interruptible Agents in Long-Horizon Web Navigation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zou, Henry Peng, Miao, Chunyu, Huang, Wei-Chieh, Chen, Yankai, Zhou, Yue, Zhang, Hanrong, Wu, Yaozu, Fang, Liancheng, Gu, Zhengyao, Zhang, Zhen, Zheng, Kening, Wang, Fangxin, Nian, Yi, Li, Shanghao, Fan, Wenzhe, He, Langzhou, Zhang, Weizhi, Liu, Xue, Yu, Philip S. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
por: Zou, Henry Peng, et al.
Publicado: (2025)
por: Zou, Henry Peng, et al.
Publicado: (2025)
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
por: Huang, Wei-Chieh, et al.
Publicado: (2025)
por: Huang, Wei-Chieh, et al.
Publicado: (2025)
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
por: Zou, Henry Peng, et al.
Publicado: (2025)
por: Zou, Henry Peng, et al.
Publicado: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
por: Gu, Zhengyao, et al.
Publicado: (2025)
por: Gu, Zhengyao, et al.
Publicado: (2025)
TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency
por: Zou, Henry Peng, et al.
Publicado: (2025)
por: Zou, Henry Peng, et al.
Publicado: (2025)
Embracing Trustworthy Brain-Agent Collaboration as Paradigm Extension for Intelligent Assistive Technologies
por: Chen, Yankai, et al.
Publicado: (2025)
por: Chen, Yankai, et al.
Publicado: (2025)
TabGen-ICL: Residual-Aware In-Context Example Selection for Tabular Data Generation
por: Fang, Liancheng, et al.
Publicado: (2025)
por: Fang, Liancheng, et al.
Publicado: (2025)
Locally Confident, Globally Stuck: The Quality-Exploration Dilemma in Diffusion Language Models
por: Fang, Liancheng, et al.
Publicado: (2026)
por: Fang, Liancheng, et al.
Publicado: (2026)
Multi-Agent Autonomous Driving Systems with Large Language Models: A Survey of Recent Advances
por: Wu, Yaozu, et al.
Publicado: (2025)
por: Wu, Yaozu, et al.
Publicado: (2025)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
por: Wang, Fangxin, et al.
Publicado: (2026)
por: Wang, Fangxin, et al.
Publicado: (2026)
TodyComm: Task-Oriented Dynamic Communication for Multi-Round LLM-based Multi-Agent System
por: Fan, Wenzhe, et al.
Publicado: (2026)
por: Fan, Wenzhe, et al.
Publicado: (2026)
PSG-Agent: Personality-Aware Safety Guardrail for LLM-based Agents
por: Wu, Yaozu, et al.
Publicado: (2025)
por: Wu, Yaozu, et al.
Publicado: (2025)
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
por: Zhang, Hanrong, et al.
Publicado: (2026)
por: Zhang, Hanrong, et al.
Publicado: (2026)
RECODE-H: A Benchmark for Research Code Development with Interactive Human Feedback
por: Miao, Chunyu, et al.
Publicado: (2025)
por: Miao, Chunyu, et al.
Publicado: (2025)
MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample Selection
por: Fang, Liancheng, et al.
Publicado: (2025)
por: Fang, Liancheng, et al.
Publicado: (2025)
Do We Really Need Graph Convolution During Training? Light Post-Training Graph-ODE for Efficient Recommendation
por: Zhang, Weizhi, et al.
Publicado: (2024)
por: Zhang, Weizhi, et al.
Publicado: (2024)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
por: He, Langzhou, et al.
Publicado: (2026)
por: He, Langzhou, et al.
Publicado: (2026)
GAM: Hierarchical Graph-based Agentic Memory for LLM Agents
por: Wu, Zhaofen, et al.
Publicado: (2026)
por: Wu, Zhaofen, et al.
Publicado: (2026)
Unveiling Language Routing Isolation in Multilingual MoE Models for Interpretable Subnetwork Adaptation
por: Zheng, Kening, et al.
Publicado: (2026)
por: Zheng, Kening, et al.
Publicado: (2026)
From Web Search towards Agentic Deep Research: Incentivizing Search with Reasoning Agents
por: Zhang, Weizhi, et al.
Publicado: (2025)
por: Zhang, Weizhi, et al.
Publicado: (2025)
Distributionally Robust Set Representation Learning Under Inference-Time Element Corruption
por: Chen, Yankai, et al.
Publicado: (2026)
por: Chen, Yankai, et al.
Publicado: (2026)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
por: Zou, Henry Peng, et al.
Publicado: (2024)
por: Zou, Henry Peng, et al.
Publicado: (2024)
DarkMind: Latent Chain-of-Thought Backdoor in Customized LLMs
por: Guo, Zhen, et al.
Publicado: (2025)
por: Guo, Zhen, et al.
Publicado: (2025)
ARC-STAR: Auditable Post-Hoc Correction for PDE Foundation Models
por: Li, Chengze, et al.
Publicado: (2026)
por: Li, Chengze, et al.
Publicado: (2026)
Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Training
por: Gu, Zhengyao, et al.
Publicado: (2026)
por: Gu, Zhengyao, et al.
Publicado: (2026)
From Uncertainty to Clarity: Uncertainty-Guided Class-Incremental Learning for Limited Biomedical Samples via Semantic Expansion
por: Yao, Yifei, et al.
Publicado: (2024)
por: Yao, Yifei, et al.
Publicado: (2024)
Enabling Deterministic User-Level Interrupts in Real-Time Processors via Hardware Extension
por: Yang, Hongbin, et al.
Publicado: (2026)
por: Yang, Hongbin, et al.
Publicado: (2026)
MANSY: Generalizing Neural Adaptive Immersive Video Streaming With Ensemble and Representation Learning
por: Wu, Duo, et al.
Publicado: (2023)
por: Wu, Duo, et al.
Publicado: (2023)
MemoryCD: Benchmarking Long-Context User Memory of LLM Agents for Lifelong Cross-Domain Personalization
por: Zhang, Weizhi, et al.
Publicado: (2026)
por: Zhang, Weizhi, et al.
Publicado: (2026)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
por: Zhang, Lingzhe, et al.
Publicado: (2026)
por: Zhang, Lingzhe, et al.
Publicado: (2026)
Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
por: Li, Shanghao, et al.
Publicado: (2026)
por: Li, Shanghao, et al.
Publicado: (2026)
Detecting Hallucinations in Graph Retrieval-Augmented Generation via Attention Patterns and Semantic Alignment
por: Li, Shanghao, et al.
Publicado: (2025)
por: Li, Shanghao, et al.
Publicado: (2025)
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
por: Zhao, Bingchen, et al.
Publicado: (2026)
por: Zhao, Bingchen, et al.
Publicado: (2026)
Handling The Non-Smooth Challenge in Tensor SVD: A Multi-Objective Tensor Recovery Framework
por: Zheng, Jingjing, et al.
Publicado: (2023)
por: Zheng, Jingjing, et al.
Publicado: (2023)
Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization
por: He, Langzhou, et al.
Publicado: (2025)
por: He, Langzhou, et al.
Publicado: (2025)
LawMind: A Law-Driven Paradigm for Discovering Analytical Solutions to Partial Differential Equations
por: Zheng, Min-Yi, et al.
Publicado: (2026)
por: Zheng, Min-Yi, et al.
Publicado: (2026)
AdaptiveK: Complexity-Driven Sparse Autoencoders for Interpretable Language Model Representations
por: Yao, Yifei, et al.
Publicado: (2025)
por: Yao, Yifei, et al.
Publicado: (2025)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
por: Li, Yangning, et al.
Publicado: (2025)
por: Li, Yangning, et al.
Publicado: (2025)
TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models
por: Zhang, Liancheng, et al.
Publicado: (2026)
por: Zhang, Liancheng, et al.
Publicado: (2026)
Tough and Recyclable Poly(phthalazinone ether sulfone ketone) Aerogel Fibers
por: Chen Zhuo, et al.
Publicado: (2026)
por: Chen Zhuo, et al.
Publicado: (2026)
Ejemplares similares
-
LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey
por: Zou, Henry Peng, et al.
Publicado: (2025) -
Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
por: Huang, Wei-Chieh, et al.
Publicado: (2025) -
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
por: Zou, Henry Peng, et al.
Publicado: (2025) -
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
por: Gu, Zhengyao, et al.
Publicado: (2025) -
TestNUC: Enhancing Test-Time Computing Approaches and Scaling through Neighboring Unlabeled Data Consistency
por: Zou, Henry Peng, et al.
Publicado: (2025)