(Computer) Vision in Action: Comparing Remote Sighted Assistance and a Multimodal Voice Agent in Inspection Sequences
Fuente:
arXiv
Guardado en:
| Autores principales: | Rudaz, Damien, Carreras, Barbara Nino, Merlino, Sara, Due, Brian L., Brown, Barry |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Everything Counts: The Managed Omnirelevance of Speech in Human-Voice Agent Interaction
por: Rudaz, Damien, et al.
Publicado: (2025)
por: Rudaz, Damien, et al.
Publicado: (2025)
"Playing the robot's advocate": Bystanders' descriptions of a robot's conduct in public settings
por: Rudaz, Damien, et al.
Publicado: (2025)
por: Rudaz, Damien, et al.
Publicado: (2025)
Public speech recognition transcripts as a configuring parameter
por: Rudaz, Damien, et al.
Publicado: (2025)
por: Rudaz, Damien, et al.
Publicado: (2025)
Sharing Construction Safety Inspection Experiences and Site-Specific Knowledge through XR-Augmented Visual Assistance
por: Liu, Pengkun, et al.
Publicado: (2022)
por: Liu, Pengkun, et al.
Publicado: (2022)
Utilizing Vision-Language Models as Action Models for Intent Recognition and Assistance
por: Contreras, Cesar Alan, et al.
Publicado: (2025)
por: Contreras, Cesar Alan, et al.
Publicado: (2025)
Immutable Explainability: Fuzzy Logic and Blockchain for Verifiable Affective AI
por: Fransoy, Marcelo, et al.
Publicado: (2025)
por: Fransoy, Marcelo, et al.
Publicado: (2025)
MarkupLens: Balancing Computer Vision Assistance and Control in Professional Video Annotation for Video-Based Design Tasks
por: He, Tianhao, et al.
Publicado: (2024)
por: He, Tianhao, et al.
Publicado: (2024)
Human-Centered Design and Evaluation of a Workplace for the Remote Assistance of Highly Automated Vehicles
por: Schrank, Andreas, et al.
Publicado: (2023)
por: Schrank, Andreas, et al.
Publicado: (2023)
Multimodal Programming in Computer Science with Interactive Assistance Powered by Large Language Model
por: Gupta, Rajan Das, et al.
Publicado: (2025)
por: Gupta, Rajan Das, et al.
Publicado: (2025)
Enabling Additive Manufacturing Part Inspection of Digital Twins via Collaborative Virtual Reality
por: Chheang, Vuthea, et al.
Publicado: (2024)
por: Chheang, Vuthea, et al.
Publicado: (2024)
Investigating Remote Hands-On Assistance for Collaborative Development of Embedded Systems
por: Chen, Yan, et al.
Publicado: (2024)
por: Chen, Yan, et al.
Publicado: (2024)
In-Browser Agents for Search Assistance
por: Zerhoudi, Saber, et al.
Publicado: (2026)
por: Zerhoudi, Saber, et al.
Publicado: (2026)
GazePointAR: A Context-Aware Multimodal Voice Assistant for Pronoun Disambiguation in Wearable Augmented Reality
por: Lee, Jaewook, et al.
Publicado: (2024)
por: Lee, Jaewook, et al.
Publicado: (2024)
SightGlow: A Web Extension to Enhance Color Perception and Interaction for Vision Deficiency
por: Paudel, Sansrit
Publicado: (2024)
por: Paudel, Sansrit
Publicado: (2024)
Comparing Human Oversight Strategies for Computer-Use Agents
por: Chen, Chaoran, et al.
Publicado: (2026)
por: Chen, Chaoran, et al.
Publicado: (2026)
ProAgentBench: Evaluating LLM Agents for Proactive Assistance with Real-World Data
por: Tang, Yuanbo, et al.
Publicado: (2026)
por: Tang, Yuanbo, et al.
Publicado: (2026)
NoteFlow: Recommending Charts as Sight Glasses for Tracing Data Flow in Computational Notebooks
por: Tian, Yuan, et al.
Publicado: (2025)
por: Tian, Yuan, et al.
Publicado: (2025)
Leveraging GCN-based Action Recognition for Teleoperation in Daily Activity Assistance
por: Kwok, Thomas M., et al.
Publicado: (2025)
por: Kwok, Thomas M., et al.
Publicado: (2025)
A User-Centered Teleoperation GUI for Automated Vehicles: Identifying and Evaluating Information Requirements for Remote Driving and Assistance
por: Wolf, Maria-Magdalena, et al.
Publicado: (2025)
por: Wolf, Maria-Magdalena, et al.
Publicado: (2025)
Multi-Sensor Fusion-Based Mobile Manipulator Remote Control for Intelligent Smart Home Assistance
por: Jin, Xiao, et al.
Publicado: (2025)
por: Jin, Xiao, et al.
Publicado: (2025)
ToPSen: Task-Oriented Priming and Sensory Alignment for Comparing Coding Strategies Between Sighted and Blind Programmers
por: Ehtesham-Ul-Haque, Md, et al.
Publicado: (2025)
por: Ehtesham-Ul-Haque, Md, et al.
Publicado: (2025)
Emerging Practices for Large Multimodal Model (LMM) Assistance for People with Visual Impairments: Implications for Design
por: Xie, Jingyi, et al.
Publicado: (2024)
por: Xie, Jingyi, et al.
Publicado: (2024)
Generative Muscle Stimulation: Providing Users with Physical Assistance by Constraining Multimodal-AI with Embodied Knowledge
por: Ho, Yun, et al.
Publicado: (2025)
por: Ho, Yun, et al.
Publicado: (2025)
Beyond Chat and Clicks: GUI Agents for In-Situ Assistance via Live Interface Transformation
por: Hao, Pan, et al.
Publicado: (2026)
por: Hao, Pan, et al.
Publicado: (2026)
GAIPAT -Dataset on Human Gaze and Actions for Intent Prediction in Assembly Tasks
por: Grand, Maxence, et al.
Publicado: (2025)
por: Grand, Maxence, et al.
Publicado: (2025)
TermSight: Making Service Contracts Approachable
por: Huang, Ziheng, et al.
Publicado: (2025)
por: Huang, Ziheng, et al.
Publicado: (2025)
How Voice and Helpfulness Shape Perceptions in Human-Agent Teams
por: Westby, Samuel, et al.
Publicado: (2023)
por: Westby, Samuel, et al.
Publicado: (2023)
Comparing Continuous and Retrospective Emotion Ratings in Remote VR Study
por: Warsinke, Maximilian, et al.
Publicado: (2024)
por: Warsinke, Maximilian, et al.
Publicado: (2024)
TUM Teleoperation: Open Source Software for Remote Driving and Assistance of Automated Vehicles
por: Kerbl, Tobias, et al.
Publicado: (2025)
por: Kerbl, Tobias, et al.
Publicado: (2025)
Digital Eyes: Social Implications of XR EyeSight
por: Vergari, Maurizio, et al.
Publicado: (2024)
por: Vergari, Maurizio, et al.
Publicado: (2024)
Can AI Prompt Humans? Multimodal Agents Prompt Players' Game Actions and Show Consequences to Raise Sustainability Awareness
por: Zhang, Qinshi, et al.
Publicado: (2024)
por: Zhang, Qinshi, et al.
Publicado: (2024)
Exploring the Interplay Between Voice, Personality, and Gender in Human-Agent Interactions
por: Hackney, Kai Alexander, et al.
Publicado: (2026)
por: Hackney, Kai Alexander, et al.
Publicado: (2026)
LLM-Mediated Domain-Specific Voice Agents: The Case of TextileBot
por: Zhong, Shu, et al.
Publicado: (2024)
por: Zhong, Shu, et al.
Publicado: (2024)
Human-Augmented Reality Interaction in Rebar Inspection
por: Sanei, Mahsa, et al.
Publicado: (2026)
por: Sanei, Mahsa, et al.
Publicado: (2026)
Making Videos Accessible for Blind and Low Vision Users Using a Multimodal Agent Video Player
por: Olmos, Adriana, et al.
Publicado: (2026)
por: Olmos, Adriana, et al.
Publicado: (2026)
Immersive Analysis: Enhancing Material Inspection of X-Ray Computed Tomography Datasets in Augmented Reality
por: Gall, Alexander, et al.
Publicado: (2024)
por: Gall, Alexander, et al.
Publicado: (2024)
More AI Assistance Reduces Cognitive Engagement: Examining the AI Assistance Dilemma in AI-Supported Note-Taking
por: Chen, Xinyue, et al.
Publicado: (2025)
por: Chen, Xinyue, et al.
Publicado: (2025)
"The Guide Has Your Back": Exploring How Sighted Guides Can Enhance Accessibility in Social Virtual Reality for Blind and Low Vision People
por: Collins, Jazmin, et al.
Publicado: (2024)
por: Collins, Jazmin, et al.
Publicado: (2024)
Voice to Vision: Enhancing Civic Decision-Making through Co-Designed Data Infrastructure
por: Hughes, Maggie, et al.
Publicado: (2025)
por: Hughes, Maggie, et al.
Publicado: (2025)
Designing for Adolescent Voice in Health Decisions: Embodied Conversational Agents for HPV Vaccination
por: Steenstra, Ian, et al.
Publicado: (2026)
por: Steenstra, Ian, et al.
Publicado: (2026)
Ejemplares similares
-
Everything Counts: The Managed Omnirelevance of Speech in Human-Voice Agent Interaction
por: Rudaz, Damien, et al.
Publicado: (2025) -
"Playing the robot's advocate": Bystanders' descriptions of a robot's conduct in public settings
por: Rudaz, Damien, et al.
Publicado: (2025) -
Public speech recognition transcripts as a configuring parameter
por: Rudaz, Damien, et al.
Publicado: (2025) -
Sharing Construction Safety Inspection Experiences and Site-Specific Knowledge through XR-Augmented Visual Assistance
por: Liu, Pengkun, et al.
Publicado: (2022) -
Utilizing Vision-Language Models as Action Models for Intent Recognition and Assistance
por: Contreras, Cesar Alan, et al.
Publicado: (2025)