VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhao, Baoquan, Ma, Xiaofan, Pang, Qianshi, Wang, Ruomei, Zhou, Fan, Lin, Shujin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
por: Ghosh, Indrajeet, et al.
Publicado: (2025)
por: Ghosh, Indrajeet, et al.
Publicado: (2025)
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
por: Wang, Yingna, et al.
Publicado: (2025)
por: Wang, Yingna, et al.
Publicado: (2025)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
por: Wong, Kam Kwai, et al.
Publicado: (2023)
por: Wong, Kam Kwai, et al.
Publicado: (2023)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
por: Ning, Zheng, et al.
Publicado: (2024)
por: Ning, Zheng, et al.
Publicado: (2024)
Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions
por: Teskeredzic, Edvin, et al.
Publicado: (2025)
por: Teskeredzic, Edvin, et al.
Publicado: (2025)
Unveiling the Visual Rhetoric of Persuasive Cartography: A Case Study of the Design of Octopus Maps
por: Lin, Daocheng, et al.
Publicado: (2025)
por: Lin, Daocheng, et al.
Publicado: (2025)
DIGITWISE: Digital Twin-based Modeling of Adaptive Video Streaming Engagement
por: Artioli, Emanuele, et al.
Publicado: (2025)
por: Artioli, Emanuele, et al.
Publicado: (2025)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
por: Shi, Weiyan, et al.
Publicado: (2026)
por: Shi, Weiyan, et al.
Publicado: (2026)
M2AR: A Web-based Modeling Environment for the Augmented Reality Workflow Modeling Language
por: Muff, Fabian, et al.
Publicado: (2024)
por: Muff, Fabian, et al.
Publicado: (2024)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
por: Fu, Chencan, et al.
Publicado: (2024)
por: Fu, Chencan, et al.
Publicado: (2024)
Hue4U: Real-Time Personalized Color Correction in Augmented Reality
por: Qin, Jingwen, et al.
Publicado: (2025)
por: Qin, Jingwen, et al.
Publicado: (2025)
Development of Immersive Virtual and Augmented Reality-Based Joint Attention Training Platform for Children with Autism
por: Samantaray, Ashirbad, et al.
Publicado: (2025)
por: Samantaray, Ashirbad, et al.
Publicado: (2025)
MV-Crafter: An Intelligent System for Music-guided Video Generation
por: Chen, Chuer, et al.
Publicado: (2025)
por: Chen, Chuer, et al.
Publicado: (2025)
Winds Through Time: Interactive Data Visualization and Physicalization for Paleoclimate Communication
por: Hunter, David, et al.
Publicado: (2025)
por: Hunter, David, et al.
Publicado: (2025)
MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
por: Ning, Zheng, et al.
Publicado: (2024)
por: Ning, Zheng, et al.
Publicado: (2024)
Save It for the "Hot" Day: An LLM-Empowered Visual Analytics System for Heat Risk Management
por: Li, Haobo, et al.
Publicado: (2024)
por: Li, Haobo, et al.
Publicado: (2024)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
por: Wang, Xianghan
Publicado: (2025)
por: Wang, Xianghan
Publicado: (2025)
Visual instrument co-design embracing the unique movement capabilities of a dancer with physical disability
por: Trolland, Sam, et al.
Publicado: (2024)
por: Trolland, Sam, et al.
Publicado: (2024)
Comparing Visual Metaphors with Textual Code For Learning Basic Computer Science Concepts in Virtual Reality
por: Baron, Kevin William
Publicado: (2024)
por: Baron, Kevin William
Publicado: (2024)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
por: Fan, Hector, et al.
Publicado: (2026)
por: Fan, Hector, et al.
Publicado: (2026)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025)
por: Shi, Weiyan, et al.
Publicado: (2025)
Computational Analysis of Stress, Depression and Engagement in Mental Health: A Survey
por: Kumar, Puneet, et al.
Publicado: (2024)
por: Kumar, Puneet, et al.
Publicado: (2024)
Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
por: He, Xu, et al.
Publicado: (2024)
por: He, Xu, et al.
Publicado: (2024)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
por: Song, Jiarun, et al.
Publicado: (2026)
por: Song, Jiarun, et al.
Publicado: (2026)
ICE: Interactive 3D Game Character Editing via Dialogue
por: Wu, Haoqian, et al.
Publicado: (2024)
por: Wu, Haoqian, et al.
Publicado: (2024)
G3R: Generating Rich and Fine-grained mmWave Radar Data from 2D Videos for Generalized Gesture Recognition
por: Deng, Kaikai, et al.
Publicado: (2024)
por: Deng, Kaikai, et al.
Publicado: (2024)
Adaptive Virtual Reality Museum: A Closed-Loop Framewor for Engagement-Aware Cultural Heritage
por: Damouni, Joseph, et al.
Publicado: (2026)
por: Damouni, Joseph, et al.
Publicado: (2026)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
por: De La Fuente, Tessa
Publicado: (2025)
por: De La Fuente, Tessa
Publicado: (2025)
ScaleTrotter: Illustrative Visual Travels Across Negative Scales
por: Halladjian, Sarkis, et al.
Publicado: (2019)
por: Halladjian, Sarkis, et al.
Publicado: (2019)
Multimodal Infusion Tuning for Large Models
por: Sun, Hao, et al.
Publicado: (2024)
por: Sun, Hao, et al.
Publicado: (2024)
Impact Ambivalence: How People with Eating Disorders Get Trapped in the Perpetual Cycle of Digital Food Content Engagement
por: Choi, Ryuhaerang, et al.
Publicado: (2023)
por: Choi, Ryuhaerang, et al.
Publicado: (2023)
MetaDragonBoat: Exploring Paddling Techniques of Virtual Dragon Boating in a Metaverse Campus
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
A Collaborative Extended Reality Prototype for 3D Surgical Planning and Visualization
por: Qiu, Shi, et al.
Publicado: (2026)
por: Qiu, Shi, et al.
Publicado: (2026)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
por: Zhou, Dongliang, et al.
Publicado: (2025)
por: Zhou, Dongliang, et al.
Publicado: (2025)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
por: Liu, Qixuan, et al.
Publicado: (2025)
por: Liu, Qixuan, et al.
Publicado: (2025)
T2VTree: User-Centered Visual Analytics for Agent-Assisted Thought-to-Video Authoring
por: Zheng, Zhuoyun, et al.
Publicado: (2026)
por: Zheng, Zhuoyun, et al.
Publicado: (2026)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
por: Qiu, Yue, et al.
Publicado: (2025)
por: Qiu, Yue, et al.
Publicado: (2025)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
por: Yeh, Catherine, et al.
Publicado: (2026)
por: Yeh, Catherine, et al.
Publicado: (2026)
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
por: Wang, Bryan, et al.
Publicado: (2024)
por: Wang, Bryan, et al.
Publicado: (2024)
Ejemplares similares
-
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
por: Ghosh, Indrajeet, et al.
Publicado: (2025) -
Facilitating Daily Practice in Intangible Cultural Heritage through Virtual Reality: A Case Study of Traditional Chinese Flower Arrangement
por: Wang, Yingna, et al.
Publicado: (2025) -
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
por: Wong, Kam Kwai, et al.
Publicado: (2023) -
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
por: Ning, Zheng, et al.
Publicado: (2024) -
Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions
por: Teskeredzic, Edvin, et al.
Publicado: (2025)