VideoA11y: Method and Dataset for Accessible Video Description
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Chaoyu, Padmanabhuni, Sid, Cheema, Maryam, Seifi, Hasti, Fazli, Pooyan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
por: Cheema, Maryam, et al.
Publicado: (2024)
por: Cheema, Maryam, et al.
Publicado: (2024)
ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
por: Cheema, Maryam, et al.
Publicado: (2026)
por: Cheema, Maryam, et al.
Publicado: (2026)
DescribePro: Collaborative Audio Description with Human-AI Interaction
por: Cheema, Maryam, et al.
Publicado: (2025)
por: Cheema, Maryam, et al.
Publicado: (2025)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
por: Hegde, Shamanthak, et al.
Publicado: (2025)
por: Hegde, Shamanthak, et al.
Publicado: (2025)
EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
por: Leng, Zikang, et al.
Publicado: (2026)
por: Leng, Zikang, et al.
Publicado: (2026)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
por: Li, Chaoyu, et al.
Publicado: (2024)
por: Li, Chaoyu, et al.
Publicado: (2024)
VideoMix: Aggregating How-To Videos for Task-Oriented Learning
por: Yang, Saelyne, et al.
Publicado: (2025)
por: Yang, Saelyne, et al.
Publicado: (2025)
Accessible, At-Home Detection of Parkinson's Disease via Multi-task Video Analysis
por: Islam, Md Saiful, et al.
Publicado: (2024)
por: Islam, Md Saiful, et al.
Publicado: (2024)
Panonut360: A Head and Eye Tracking Dataset for Panoramic Video
por: Xu, Yutong, et al.
Publicado: (2024)
por: Xu, Yutong, et al.
Publicado: (2024)
The Visual Experience Dataset: Over 200 Recorded Hours of Integrated Eye Movement, Odometry, and Egocentric Video
por: Greene, Michelle R., et al.
Publicado: (2024)
por: Greene, Michelle R., et al.
Publicado: (2024)
A Comparison of Bounding Box and Landmark Detection Methods for Video-Based Heart Rate Estimation
por: Liang, Laurence
Publicado: (2023)
por: Liang, Laurence
Publicado: (2023)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
por: Foteinopoulou, Niki Maria, et al.
Publicado: (2023)
por: Foteinopoulou, Niki Maria, et al.
Publicado: (2023)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
por: Garg, Mallika, et al.
Publicado: (2024)
por: Garg, Mallika, et al.
Publicado: (2024)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
Steering Generative Models for Accessibility: EasyRead Image Generation
por: Dickenmann, Nicolas, et al.
Publicado: (2026)
por: Dickenmann, Nicolas, et al.
Publicado: (2026)
The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility
por: Zhang, Xiantao
Publicado: (2025)
por: Zhang, Xiantao
Publicado: (2025)
SpriteHand: Real-Time Versatile Hand-Object Interaction with Autoregressive Video Generation
por: Li, Zisu, et al.
Publicado: (2025)
por: Li, Zisu, et al.
Publicado: (2025)
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
por: Lin, David Chuan-En, et al.
Publicado: (2022)
por: Lin, David Chuan-En, et al.
Publicado: (2022)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
por: Nadeem, Asmar, et al.
Publicado: (2024)
por: Nadeem, Asmar, et al.
Publicado: (2024)
Reframe Anything: LLM Agent for Open World Video Reframing
por: Cao, Jiawang, et al.
Publicado: (2024)
por: Cao, Jiawang, et al.
Publicado: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
por: Huh, Mina, et al.
Publicado: (2025)
por: Huh, Mina, et al.
Publicado: (2025)
Analyzing Swimming Performance Using Drone Captured Aerial Videos
por: Tran, Thu, et al.
Publicado: (2025)
por: Tran, Thu, et al.
Publicado: (2025)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
por: Eing, Lennart, et al.
Publicado: (2026)
por: Eing, Lennart, et al.
Publicado: (2026)
ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
por: He, Yuchen, et al.
Publicado: (2025)
por: He, Yuchen, et al.
Publicado: (2025)
Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
por: Zhou, Puqi, et al.
Publicado: (2026)
por: Zhou, Puqi, et al.
Publicado: (2026)
Hybrid 3D Human Pose Estimation with Monocular Video and Sparse IMUs
por: Bao, Yiming, et al.
Publicado: (2024)
por: Bao, Yiming, et al.
Publicado: (2024)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
por: Li, Chaoyu, et al.
Publicado: (2025)
por: Li, Chaoyu, et al.
Publicado: (2025)
Multiscaled Multi-Head Attention-based Video Transformer Network for Hand Gesture Recognition
por: Garg, Mallika, et al.
Publicado: (2025)
por: Garg, Mallika, et al.
Publicado: (2025)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
por: Zheng, Ruizhe, et al.
Publicado: (2024)
por: Zheng, Ruizhe, et al.
Publicado: (2024)
CinePreGen: Camera Controllable Video Previsualization via Engine-powered Diffusion
por: Chen, Yiran, et al.
Publicado: (2024)
por: Chen, Yiran, et al.
Publicado: (2024)
Acoustic Field Video for Multimodal Scene Understanding
por: Kim, Daehwa, et al.
Publicado: (2026)
por: Kim, Daehwa, et al.
Publicado: (2026)
Detecting Activities of Daily Living in Egocentric Video to Contextualize Hand Use at Home in Outpatient Neurorehabilitation Settings
por: Kadambi, Adesh, et al.
Publicado: (2024)
por: Kadambi, Adesh, et al.
Publicado: (2024)
Surgment: Segmentation-enabled Semantic Search and Creation of Visual Question and Feedback to Support Video-Based Surgery Learning
por: Wang, Jingying, et al.
Publicado: (2024)
por: Wang, Jingying, et al.
Publicado: (2024)
SVFAP: Self-supervised Video Facial Affect Perceiver
por: Sun, Licai, et al.
Publicado: (2023)
por: Sun, Licai, et al.
Publicado: (2023)
Editing Physiological Signals in Videos Using Latent Representations
por: Zhou, Tianwen, et al.
Publicado: (2025)
por: Zhou, Tianwen, et al.
Publicado: (2025)
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
por: Kulkarni, Yogesh, et al.
Publicado: (2025)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024)
por: He, Xu, et al.
Publicado: (2024)
AdapTics: A Toolkit for Creative Design and Integration of Real-Time Adaptive Mid-Air Ultrasound Tactons
por: John, Kevin, et al.
Publicado: (2024)
por: John, Kevin, et al.
Publicado: (2024)
Ejemplares similares
-
Describe Now: User-Driven Audio Description for Blind and Low Vision Individuals
por: Cheema, Maryam, et al.
Publicado: (2024) -
ViDscribe: Multimodal AI for Customizing Audio Description and Question Answering in Online Videos
por: Cheema, Maryam, et al.
Publicado: (2026) -
DescribePro: Collaborative Audio Description with Human-AI Interaction
por: Cheema, Maryam, et al.
Publicado: (2025) -
ChartQA-X: Generating Explanations for Visual Chart Reasoning
por: Hegde, Shamanthak, et al.
Publicado: (2025) -
EduGage: Methods and Dataset for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
por: Leng, Zikang, et al.
Publicado: (2026)