Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yayuan, Li, Chenglin, Wang, Jingying, Bellos, Filippos, Guo, Anhong, Corso, Jason J. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Consistent Long-Term Pose Generation
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
WorldScribe: Towards Context-Aware Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
Rubikon: Intelligent Tutoring for Rubik's Cube Learning Through AR-enabled Physical Task Reconfiguration
by: Ren, Haocheng, et al.
Published: (2025)
by: Ren, Haocheng, et al.
Published: (2025)
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
by: Kim, Sejin, et al.
Published: (2024)
by: Kim, Sejin, et al.
Published: (2024)
Probing the Gaps in ChatGPT Live Video Chat for Real-World Assistance for People who are Blind or Visually Impaired
by: Chang, Ruei-Che, et al.
Published: (2025)
by: Chang, Ruei-Che, et al.
Published: (2025)
HandProxy: Expanding the Affordances of Speech Interfaces in Immersive Environments with a Virtual Proxy Hand
by: Liang, Chen, et al.
Published: (2025)
by: Liang, Chen, et al.
Published: (2025)
Looking Together $\neq$ Seeing the Same Thing: Understanding Surgeons' Visual Needs During Intra-operative Coordination and Instruction
by: Popov, Vitaliy, et al.
Published: (2024)
by: Popov, Vitaliy, et al.
Published: (2024)
TouchScribe: Augmenting Non-Visual Hand-Object Interactions with Automated Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2026)
by: Chang, Ruei-Che, et al.
Published: (2026)
EditScribe: Non-Visual Image Editing with Natural Language Verification Loops
by: Chang, Ruei-Che, et al.
Published: (2024)
by: Chang, Ruei-Che, et al.
Published: (2024)
ProgramAlly: Creating Custom Visual Access Programs via Multi-Modal End-User Programming
by: Herskovitz, Jaylin, et al.
Published: (2024)
by: Herskovitz, Jaylin, et al.
Published: (2024)
Audio Description Customization
by: Natalie, Rosiana, et al.
Published: (2024)
by: Natalie, Rosiana, et al.
Published: (2024)
A11y-CUA Dataset: Characterizing the Accessibility Gap in Computer Use Agents
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
by: Mohanbabu, Ananya Gubbi, et al.
Published: (2026)
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
by: Wang, Qixin, et al.
Published: (2025)
by: Wang, Qixin, et al.
Published: (2025)
Visualizing the Invisible: A Generative AR System for Intuitive Multi-Modal Sensor Data Presentation
by: Guo, Yunqi, et al.
Published: (2024)
by: Guo, Yunqi, et al.
Published: (2024)
Beyond Screenshots: Evaluating VLMs' Understanding of UI Animations
by: Liang, Chen, et al.
Published: (2026)
by: Liang, Chen, et al.
Published: (2026)
InteractOut: Leveraging Interaction Proxies as Input Manipulation Strategies for Reducing Smartphone Overuse
by: Lu, Tao, et al.
Published: (2024)
by: Lu, Tao, et al.
Published: (2024)
Connecting Dreams with Visual Brainstorming Instruction
by: Sun, Yasheng, et al.
Published: (2024)
by: Sun, Yasheng, et al.
Published: (2024)
Making the Invisible Visible: Toward Micro-Expression Visualization for Empathy in Social Interaction
by: Yin, Feiyang, et al.
Published: (2026)
by: Yin, Feiyang, et al.
Published: (2026)
eXplainMR: Generating Real-time Textual and Visual eXplanations to Facilitate UltraSonography Learning in MR
by: Wang, Jingying, et al.
Published: (2025)
by: Wang, Jingying, et al.
Published: (2025)
Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning
by: Wang, Jingying, et al.
Published: (2024)
by: Wang, Jingying, et al.
Published: (2024)
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
by: Bo, Jessica Y., et al.
Published: (2025)
by: Bo, Jessica Y., et al.
Published: (2025)
ConceptThread: Visualizing Threaded Concepts in MOOC Videos
by: Zhou, Zhiguang, et al.
Published: (2024)
by: Zhou, Zhiguang, et al.
Published: (2024)
A Recipe for Success? Exploring Strategies for Improving Non-Visual Access to Cooking Instructions
by: Li, Franklin Mingzhe, et al.
Published: (2024)
by: Li, Franklin Mingzhe, et al.
Published: (2024)
DeckFlow: Iterative Specification on a Multimodal Generative Canvas
by: Croisdale, Gregory, et al.
Published: (2025)
by: Croisdale, Gregory, et al.
Published: (2025)
Aria-UI: Visual Grounding for GUI Instructions
by: Yang, Yuhao, et al.
Published: (2024)
by: Yang, Yuhao, et al.
Published: (2024)
Chart2Vec: A Universal Embedding of Context-Aware Visualizations
by: Chen, Qing, et al.
Published: (2023)
by: Chen, Qing, et al.
Published: (2023)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
Auditorily Embodied Conversational Agents: Effects of Spatialization and Situated Audio Cues on Presence and Social Perception
by: Cheng, Yi Fei, et al.
Published: (2026)
by: Cheng, Yi Fei, et al.
Published: (2026)
HyperMOOC: Augmenting MOOC Videos with Concept-based Embedded Visualizations
by: Ye, Li, et al.
Published: (2025)
by: Ye, Li, et al.
Published: (2025)
Insights into Natural Language Database Query Errors: From Attention Misalignment to User Handling Strategies
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Learning Password Best Practices Through In-Task Instruction
by: Ma, Qian, et al.
Published: (2026)
by: Ma, Qian, et al.
Published: (2026)
Context-KG: Context-Aware Knowledge Graph Visualization with User Preferences and Ontological Guidance
by: Perera, Rumali, et al.
Published: (2026)
by: Perera, Rumali, et al.
Published: (2026)
COIVis: Eye-tracking-based Visual Exploration of Concept Learning in MOOC Videos
by: Zhou, Zhiguang, et al.
Published: (2025)
by: Zhou, Zhiguang, et al.
Published: (2025)
LightVA: Lightweight Visual Analytics with LLM Agent-Based Task Planning and Execution
by: Zhao, Yuheng, et al.
Published: (2024)
by: Zhao, Yuheng, et al.
Published: (2024)
Turning Text and Imagery into Captivating Visual Video
by: Wang, Mingming, et al.
Published: (2024)
by: Wang, Mingming, et al.
Published: (2024)
From Awareness to Intent: Mitigating Silent Driving System Failures through Prospective Situation Awareness Enhancing Interfaces
by: Wang, Jiyao, et al.
Published: (2026)
by: Wang, Jiyao, et al.
Published: (2026)
Grid Labeling: Crowdsourcing Task-Specific Importance from Visualizations
by: Chang, Minsuk, et al.
Published: (2025)
by: Chang, Minsuk, et al.
Published: (2025)
Investigating the Task Load of Investigating the Task Load in Visualization Studies
by: Pahr, Daniel, et al.
Published: (2025)
by: Pahr, Daniel, et al.
Published: (2025)
JumpStarter: Human-AI Planning with Task-Structured Context Curation
by: Zhang, Xuanming, et al.
Published: (2024)
by: Zhang, Xuanming, et al.
Published: (2024)
Similar Items
-
Towards Consistent Long-Term Pose Generation
by: Li, Yayuan, et al.
Published: (2025) -
WorldScribe: Towards Context-Aware Live Visual Descriptions
by: Chang, Ruei-Che, et al.
Published: (2024) -
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025) -
Rubikon: Intelligent Tutoring for Rubik's Cube Learning Through AR-enabled Physical Task Reconfiguration
by: Ren, Haocheng, et al.
Published: (2025) -
Addressing and Visualizing Misalignments in Human Task-Solving Trajectories
by: Kim, Sejin, et al.
Published: (2024)