SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ning, Zheng, Wimer, Brianna L., Jiang, Kaiwen, Chen, Keyi, Ban, Jerrick, Tian, Yapeng, Zhao, Yuhang, Li, Toby Jia-Jun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
von: Ning, Zheng, et al.
Veröffentlicht: (2024)
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024)
Text-to-Audio Generation Synchronized with Videos
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Adaptive Offloading and Enhancement for Low-Light Video Analytics on Mobile Devices
von: He, Yuanyi, et al.
Veröffentlicht: (2024)
von: He, Yuanyi, et al.
Veröffentlicht: (2024)
Semantic Grouping Network for Audio Source Separation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
von: Ren, Yong, et al.
Veröffentlicht: (2025)
von: Ren, Yong, et al.
Veröffentlicht: (2025)
diveXplore 6.0: ITEC's Interactive Video Exploration System at VBS 2022
von: Leibetseder, Andreas, et al.
Veröffentlicht: (2025)
von: Leibetseder, Andreas, et al.
Veröffentlicht: (2025)
Designing a Multimodal Viewer for Piano Performance Analysis -- a Pedagogy-First Approach
von: Bae, Joonhyung, et al.
Veröffentlicht: (2025)
von: Bae, Joonhyung, et al.
Veröffentlicht: (2025)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
von: Zhao, Lingzhi, et al.
Veröffentlicht: (2025)
Think before You Leap: Content-Aware Low-Cost Edge-Assisted Video Semantic Segmentation
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
von: Yan, Mingxuan, et al.
Veröffentlicht: (2024)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Do Joint Audio-Video Generation Models Understand Physics?
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
von: Cui, Zijun, et al.
Veröffentlicht: (2026)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
Enhancing Video Music Recommendation with Transformer-Driven Audio-Visual Embeddings
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
von: Liu, Shimiao, et al.
Veröffentlicht: (2025)
MCAD: Multimodal Context-Aware Audio Description Generation For Soccer
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
von: Chaudhary, Lipisha, et al.
Veröffentlicht: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
von: Cao, Yuqin, et al.
Veröffentlicht: (2024)
von: Cao, Yuqin, et al.
Veröffentlicht: (2024)
Generalizing Video DeepFake Detection by Self-generated Audio-Visual Pseudo-Fakes
von: Wei, Zihe, et al.
Veröffentlicht: (2026)
von: Wei, Zihe, et al.
Veröffentlicht: (2026)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
MAR3: Multi-Agent Recognition, Reasoning, and Reflection for Reference Audio-Visual Segmentation
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
von: Zhao, Yuan, et al.
Veröffentlicht: (2026)
The Future is Meta: Metadata, Formats and Perspectives towards Interactive and Personalized AV Content
von: Weller, Alexander, et al.
Veröffentlicht: (2024)
von: Weller, Alexander, et al.
Veröffentlicht: (2024)
XGC-AVis: Towards Audio-Visual Content Understanding with a Multi-Agent Collaborative System
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
von: Cao, Yuqin, et al.
Veröffentlicht: (2025)
DreamFoley: Scalable VLMs for High-Fidelity Video-to-Audio Generation
von: Li, Fu, et al.
Veröffentlicht: (2025)
von: Li, Fu, et al.
Veröffentlicht: (2025)
ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2025)
von: Kawamura, Kazuki, et al.
Veröffentlicht: (2025)
Graph-based Interaction Augmentation Network for Robust Multimodal Sentiment Analysis
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
von: Zhangfeng, Hu, et al.
Veröffentlicht: (2025)
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
Camel: Frame-Level Bandwidth Estimation for Low-Latency Live Streaming under Video Bitrate Undershooting
von: Liu, Liming, et al.
Veröffentlicht: (2026)
von: Liu, Liming, et al.
Veröffentlicht: (2026)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
von: Luu, Thanh-Danh, et al.
Veröffentlicht: (2025)
Task Presentation and Human Perception in Interactive Video Retrieval
von: Willis, Nina, et al.
Veröffentlicht: (2024)
von: Willis, Nina, et al.
Veröffentlicht: (2024)
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction
von: Wang, Jiadong, et al.
Veröffentlicht: (2026)
von: Wang, Jiadong, et al.
Veröffentlicht: (2026)
Continual Audio-Visual Sound Separation
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
Building and Evaluating a Realistic Virtual World for Large Scale Urban Exploration from 360° Videos
von: Takenawa, Mizuki, et al.
Veröffentlicht: (2025)
von: Takenawa, Mizuki, et al.
Veröffentlicht: (2025)
Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
von: Yin, Shibo, et al.
Veröffentlicht: (2024)
von: Yin, Shibo, et al.
Veröffentlicht: (2024)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
von: Tsangko, Iosif, et al.
Veröffentlicht: (2026)
Interpreting Multimodal Communication at Scale in Short-Form Video: Visual, Audio, and Textual Mental Health Discourse on TikTok
von: Zha, Mingyue, et al.
Veröffentlicht: (2026)
von: Zha, Mingyue, et al.
Veröffentlicht: (2026)
Network Bending of Diffusion Models for Audio-Visual Generation
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
von: Ning, Zheng, et al.
Veröffentlicht: (2024) -
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
von: Kushwaha, Saksham Singh, et al.
Veröffentlicht: (2024) -
Text-to-Audio Generation Synchronized with Videos
von: Mo, Shentong, et al.
Veröffentlicht: (2024) -
Adaptive Offloading and Enhancement for Low-Light Video Analytics on Mobile Devices
von: He, Yuanyi, et al.
Veröffentlicht: (2024) -
Semantic Grouping Network for Audio Source Separation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)