ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Cheema, Maryam, Elahimanesh, Sina, Fazli, Pooyan, Seifi, Hasti |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DescribePro: Collaborative Audio Description with Human-AI Interaction
by: Cheema, Maryam, et al.
Published: (2025)
by: Cheema, Maryam, et al.
Published: (2025)
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
by: Cheema, Maryam, et al.
Published: (2024)
by: Cheema, Maryam, et al.
Published: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Feeds Don't Tell the Whole Story: Measuring Online-Offline Emotion Alignment
by: Elahimanesh, Sina, et al.
Published: (2026)
by: Elahimanesh, Sina, et al.
Published: (2026)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
by: Li, Yinan, et al.
Published: (2026)
by: Li, Yinan, et al.
Published: (2026)
AdapTics: A Toolkit for Creative Design and Integration of Real-Time Adaptive Mid-Air Ultrasound Tactons
by: John, Kevin, et al.
Published: (2024)
by: John, Kevin, et al.
Published: (2024)
An Interactive Tool for Simulating Mid-Air Ultrasound Tactons on the Skin
by: Lim, Chungman, et al.
Published: (2024)
by: Lim, Chungman, et al.
Published: (2024)
Designing Distinguishable Mid-Air Ultrasound Tactons with Temporal Parameters
by: Lim, Chungman, et al.
Published: (2024)
by: Lim, Chungman, et al.
Published: (2024)
Grounding Emotional Descriptions to Electrovibration Haptic Signals
by: Hu, Guimin, et al.
Published: (2024)
by: Hu, Guimin, et al.
Published: (2024)
Audio Description Customization
by: Natalie, Rosiana, et al.
Published: (2024)
by: Natalie, Rosiana, et al.
Published: (2024)
Can a Machine Feel Vibrations?: A Framework for Vibrotactile Sensation and Emotion Prediction via a Neural Network
by: Lim, Chungman, et al.
Published: (2025)
by: Lim, Chungman, et al.
Published: (2025)
Emotion Alignment: Discovering the Gap Between Social Media and Real-World Sentiments in Persian Tweets and Images
by: Elahimanesh, Sina, et al.
Published: (2025)
by: Elahimanesh, Sina, et al.
Published: (2025)
Towards Designing a Question-Answering Chatbot for Online News: Understanding Questions and Perspectives
by: Hoque, Md Naimul, et al.
Published: (2023)
by: Hoque, Md Naimul, et al.
Published: (2023)
Structure Matters: Evaluating Multi-Agents Orchestration in Generative Therapeutic Chatbots
by: Elahimanesh, Sina, et al.
Published: (2026)
by: Elahimanesh, Sina, et al.
Published: (2026)
In the Arms of a Robot: Designing Autonomous Hugging Robots with Intra-Hug Gestures
by: Block, Alexis E., et al.
Published: (2022)
by: Block, Alexis E., et al.
Published: (2022)
Sportify: Question Answering with Embedded Visualizations and Personified Narratives for Sports Video
by: Lee, Chunggi, et al.
Published: (2024)
by: Lee, Chunggi, et al.
Published: (2024)
Interaction Configurations and Prompt Guidance in Conversational AI for Question Answering in Human-AI Teams
by: Song, Jaeyoon, et al.
Published: (2025)
by: Song, Jaeyoon, et al.
Published: (2025)
AQuA: Automated Question-Answering in Software Tutorial Videos with Visual Anchors
by: Yang, Saelyne, et al.
Published: (2024)
by: Yang, Saelyne, et al.
Published: (2024)
Text Entry for XR Trove (TEXT): Collecting and Analyzing Techniques for Text Input in XR
by: Bhatia, Arpit, et al.
Published: (2025)
by: Bhatia, Arpit, et al.
Published: (2025)
From Words and Exercises to Wellness: Farsi Chatbot for Self-Attachment Technique
by: Elahimanesh, Sina, et al.
Published: (2023)
by: Elahimanesh, Sina, et al.
Published: (2023)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
by: Hegde, Shamanthak, et al.
Published: (2025)
by: Hegde, Shamanthak, et al.
Published: (2025)
GeoVisA11y: An AI-based Geovisualization Question-Answering System for Screen-Reader Users
by: Li, Chu, et al.
Published: (2026)
by: Li, Chu, et al.
Published: (2026)
Toward Human Centered Interactive Clinical Question Answering System
by: Albassam, Dina
Published: (2025)
by: Albassam, Dina
Published: (2025)
Understanding and Supporting Formal Email Exchange by Answering AI-Generated Questions
by: Miura, Yusuke, et al.
Published: (2025)
by: Miura, Yusuke, et al.
Published: (2025)
Hovering Over the Key to Text Input in XR
by: Gonzalez-Franco, Mar, et al.
Published: (2024)
by: Gonzalez-Franco, Mar, et al.
Published: (2024)
OmniQuery: Contextually Augmenting Captured Multimodal Memory to Enable Personal Question Answering
by: Li, Jiahao Nick, et al.
Published: (2024)
by: Li, Jiahao Nick, et al.
Published: (2024)
Open-Ended Multi-Modal Relational Reasoning for Video Question Answering
by: Luo, Haozheng, et al.
Published: (2020)
by: Luo, Haozheng, et al.
Published: (2020)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
SensorChat: Answering Qualitative and Quantitative Questions during Long-Term Multimodal Sensor Interactions
by: Yu, Xiaofan, et al.
Published: (2025)
by: Yu, Xiaofan, et al.
Published: (2025)
Physiological and Behavioral Modeling of Stress and Cognitive Load in Web-Based Question Answering
by: Liu, Ailin, et al.
Published: (2026)
by: Liu, Ailin, et al.
Published: (2026)
Designing and Evaluating Chain-of-Hints for Scientific Question Answering
by: Jangra, Anubhav, et al.
Published: (2025)
by: Jangra, Anubhav, et al.
Published: (2025)
Making AI Drafts Count: A Quality Threshold in Audio Description Workflows
by: Do, Lana, et al.
Published: (2026)
by: Do, Lana, et al.
Published: (2026)
Multi-Hop Question Answering: When Can Humans Help, and Where do They Struggle?
by: Su, Jinyan, et al.
Published: (2025)
by: Su, Jinyan, et al.
Published: (2025)
MYCloth: Towards Intelligent and Interactive Online T-Shirt Customization based on User's Preference
by: Liu, Yexin, et al.
Published: (2024)
by: Liu, Yexin, et al.
Published: (2024)
ADx3: A Collaborative Workflow for High-Quality Accessible Audio Description
by: Do, Lana, et al.
Published: (2026)
by: Do, Lana, et al.
Published: (2026)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Adapting Online Customer Reviews for Blind Users: A Case Study of Restaurant Reviews
by: Sunkara, Mohan, et al.
Published: (2025)
by: Sunkara, Mohan, et al.
Published: (2025)
Navig-AI-tion: Navigation by Contextual AI and Spatial Audio
by: Lystbæk, Mathias N., et al.
Published: (2026)
by: Lystbæk, Mathias N., et al.
Published: (2026)
ClassComet: Exploring and Designing AI-generated Danmaku in Educational Videos to Enhance Online Learning
by: Ji, Zipeng, et al.
Published: (2025)
by: Ji, Zipeng, et al.
Published: (2025)
SimTube: Generating Simulated Video Comments through Multimodal AI and User Personas
by: Hung, Yu-Kai, et al.
Published: (2024)
by: Hung, Yu-Kai, et al.
Published: (2024)
Similar Items
-
DescribePro: Collaborative Audio Description with Human-AI Interaction
by: Cheema, Maryam, et al.
Published: (2025) -
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
by: Cheema, Maryam, et al.
Published: (2024) -
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025) -
Feeds Don't Tell the Whole Story: Measuring Online-Offline Emotion Alignment
by: Elahimanesh, Sina, et al.
Published: (2026) -
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
by: Li, Yinan, et al.
Published: (2026)