Pro$^2$Assist: Continuous Step-Aware Proactive Assistance with Multimodal Egocentric Perception for Long-Horizon Procedural Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Lilin, Yang, Bufang, Jiang, Siyang, Liu, Kaiwei, Hou, Kaiyuan, Fan, Yuang, Chen, Hongkai, Yan, Zhenyu, Jiang, Xiaofan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
di: Yang, Bufang, et al.
Pubblicazione: (2025)
di: Yang, Bufang, et al.
Pubblicazione: (2025)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
di: Yang, Bufang, et al.
Pubblicazione: (2025)
di: Yang, Bufang, et al.
Pubblicazione: (2025)
SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions
di: Yang, Bufang, et al.
Pubblicazione: (2024)
di: Yang, Bufang, et al.
Pubblicazione: (2024)
DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge
di: Yang, Bufang, et al.
Pubblicazione: (2024)
di: Yang, Bufang, et al.
Pubblicazione: (2024)
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
di: Yang, Bufang, et al.
Pubblicazione: (2026)
di: Yang, Bufang, et al.
Pubblicazione: (2026)
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
di: Jiang, Siyang, et al.
Pubblicazione: (2025)
di: Jiang, Siyang, et al.
Pubblicazione: (2025)
TDBench: A Benchmark for Top-Down Image Understanding with Reliability Analysis of Vision-Language Models
di: Hou, Kaiyuan, et al.
Pubblicazione: (2025)
di: Hou, Kaiyuan, et al.
Pubblicazione: (2025)
Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding
di: Xu, Lilin, et al.
Pubblicazione: (2025)
di: Xu, Lilin, et al.
Pubblicazione: (2025)
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
di: Jiang, Siyang, et al.
Pubblicazione: (2025)
di: Jiang, Siyang, et al.
Pubblicazione: (2025)
Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation
di: Guo, Yunqi, et al.
Pubblicazione: (2024)
di: Guo, Yunqi, et al.
Pubblicazione: (2024)
VIAssist: Adapting Multi-modal Large Language Models for Users with Visual Impairments
di: Yang, Bufang, et al.
Pubblicazione: (2024)
di: Yang, Bufang, et al.
Pubblicazione: (2024)
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks
di: Liu, Lulin, et al.
Pubblicazione: (2026)
di: Liu, Lulin, et al.
Pubblicazione: (2026)
PACT: Proactive Asking for Continual Task Assistance in Human-Robot Collaboration
di: He, Chengbo, et al.
Pubblicazione: (2026)
di: He, Chengbo, et al.
Pubblicazione: (2026)
DeepFeature: Iterative Context-aware Feature Generation for Wearable Biosignals
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
di: Liu, Kaiwei, et al.
Pubblicazione: (2025)
From Videos to Conversations: Egocentric Instructions for Task Assistance
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2026)
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
di: Ran, Dongchuan, et al.
Pubblicazione: (2026)
di: Ran, Dongchuan, et al.
Pubblicazione: (2026)
ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices
di: Pu, Kevin, et al.
Pubblicazione: (2025)
di: Pu, Kevin, et al.
Pubblicazione: (2025)
ProActor: Timing-Aware Reinforcement Learning for Proactive Task Scheduling Agents
di: Ding, Lei, et al.
Pubblicazione: (2026)
di: Ding, Lei, et al.
Pubblicazione: (2026)
ProSec: Fortifying Code LLMs with Proactive Security Alignment
di: Xu, Xiangzhe, et al.
Pubblicazione: (2024)
di: Xu, Xiangzhe, et al.
Pubblicazione: (2024)
ProAct: A Benchmark and Multimodal Framework for Structure-Aware Proactive Response
di: Zhu, Xiaomeng, et al.
Pubblicazione: (2026)
di: Zhu, Xiaomeng, et al.
Pubblicazione: (2026)
ProGuard: Towards Proactive Multimodal Safeguard
di: Yu, Shaohan, et al.
Pubblicazione: (2025)
di: Yu, Shaohan, et al.
Pubblicazione: (2025)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
di: Tang, Yuanbo, et al.
Pubblicazione: (2026)
di: Tang, Yuanbo, et al.
Pubblicazione: (2026)
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
di: Li, Zaijing, et al.
Pubblicazione: (2024)
di: Li, Zaijing, et al.
Pubblicazione: (2024)
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks
di: Gao, Hengjian, et al.
Pubblicazione: (2026)
di: Gao, Hengjian, et al.
Pubblicazione: (2026)
ReSPEC: A Framework for Online Multispectral Sensor Reconfiguration in Dynamic Environments
di: Liu, Yanchen, et al.
Pubblicazione: (2026)
di: Liu, Yanchen, et al.
Pubblicazione: (2026)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
di: Ye, Suyu, et al.
Pubblicazione: (2025)
di: Ye, Suyu, et al.
Pubblicazione: (2025)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
di: Ye, Rui, et al.
Pubblicazione: (2025)
di: Ye, Rui, et al.
Pubblicazione: (2025)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
di: Si, Shengyu, et al.
Pubblicazione: (2026)
di: Si, Shengyu, et al.
Pubblicazione: (2026)
Temporal Preferences in Language Models for Long-Horizon Assistance
di: Mazyaki, Ali, et al.
Pubblicazione: (2025)
di: Mazyaki, Ali, et al.
Pubblicazione: (2025)
AppAgent-Pro: A Proactive GUI Agent System for Multidomain Information Integration and User Assistance
di: Zhao, Yuyang, et al.
Pubblicazione: (2025)
di: Zhao, Yuyang, et al.
Pubblicazione: (2025)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
di: Xue, Wei, et al.
Pubblicazione: (2026)
di: Xue, Wei, et al.
Pubblicazione: (2026)
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
di: Yu, Chengjun, et al.
Pubblicazione: (2026)
di: Yu, Chengjun, et al.
Pubblicazione: (2026)
FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems
di: Zhao, Minghui, et al.
Pubblicazione: (2024)
di: Zhao, Minghui, et al.
Pubblicazione: (2024)
Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2025)
di: Aggarwal, Lavisha, et al.
Pubblicazione: (2025)
Continual Multimodal Egocentric Activity Recognition via Modality-Aware Novel Detection
di: Lim, Wonseon, et al.
Pubblicazione: (2026)
di: Lim, Wonseon, et al.
Pubblicazione: (2026)
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
di: Deng, Xiang, et al.
Pubblicazione: (2025)
di: Deng, Xiang, et al.
Pubblicazione: (2025)
SEC-bench Pro: Can Language Models Solve Long-Horizon Software Security Tasks?
di: Lee, Hwiwon, et al.
Pubblicazione: (2026)
di: Lee, Hwiwon, et al.
Pubblicazione: (2026)
Large Language Models for Single-Step and Multi-Step Flight Trajectory Prediction
di: Luo, Kaiwei, et al.
Pubblicazione: (2025)
di: Luo, Kaiwei, et al.
Pubblicazione: (2025)
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video
di: Li, Ao, et al.
Pubblicazione: (2026)
di: Li, Ao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
di: Yang, Bufang, et al.
Pubblicazione: (2025) -
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
di: Yang, Bufang, et al.
Pubblicazione: (2025) -
SocialMind: LLM-based Proactive AR Social Assistive System with Human-like Perception for In-situ Live Interactions
di: Yang, Bufang, et al.
Pubblicazione: (2024) -
DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge
di: Yang, Bufang, et al.
Pubblicazione: (2024) -
SensorPersona: An LLM-Empowered System for Continual Persona Extraction from Longitudinal Mobile Sensor Streams
di: Yang, Bufang, et al.
Pubblicazione: (2026)