MIMOSA: Human-AI Co-Creation of Computational Spatial Audio Effects on Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Ning, Zheng, Zhang, Zheng, Ban, Jerrick, Jiang, Kaiwen, Gan, Ruohong, Tian, Yapeng, Li, Toby Jia-Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Whispering Water: Materializing Human-AI Dialogue as Interactive Ripples
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025)
by: Kyaw, Alexander Htet, et al.
Published: (2025)
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024)
by: Fu, Chencan, et al.
Published: (2024)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
Computational Analysis of Stress, Depression and Engagement in Mental Health: A Survey
by: Kumar, Puneet, et al.
Published: (2024)
by: Kumar, Puneet, et al.
Published: (2024)
MV-Crafter: An Intelligent System for Music-guided Video Generation
by: Chen, Chuer, et al.
Published: (2025)
by: Chen, Chuer, et al.
Published: (2025)
Anchorage: Visual Analysis of Satisfaction in Customer Service Videos via Anchor Events
by: Wong, Kam Kwai, et al.
Published: (2023)
by: Wong, Kam Kwai, et al.
Published: (2023)
PodReels: Human-AI Co-Creation of Video Podcast Teasers
by: Wang, Sitong, et al.
Published: (2023)
by: Wang, Sitong, et al.
Published: (2023)
PersoNo: Personalised Notification Urgency Classifier in Mixed Reality
by: Zheng, Jingyao, et al.
Published: (2025)
by: Zheng, Jingyao, et al.
Published: (2025)
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
by: Zhao, Baoquan, et al.
Published: (2025)
by: Zhao, Baoquan, et al.
Published: (2025)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Comparing Visual Metaphors with Textual Code For Learning Basic Computer Science Concepts in Virtual Reality
by: Baron, Kevin William
Published: (2024)
by: Baron, Kevin William
Published: (2024)
Thief of Truth: VR comics about the relationship between AI and humans
by: Bae, Joonhyung
Published: (2025)
by: Bae, Joonhyung
Published: (2025)
Physical-aware Cross-modal Adversarial Network for Wearable Sensor-based Human Action Recognition
by: Ni, Jianyuan, et al.
Published: (2023)
by: Ni, Jianyuan, et al.
Published: (2023)
Revival: Collaborative Artistic Creation through Human-AI Interactions in Musical Creativity
by: Lee, Keon Ju M., et al.
Published: (2025)
by: Lee, Keon Ju M., et al.
Published: (2025)
Human-Machine Collaboration-Guided Space Design: Combination of Machine Learning Models and Humanistic Design Concepts
by: Yang, Yuxuan
Published: (2025)
by: Yang, Yuxuan
Published: (2025)
PoEmotion: Can AI Utilize Chinese Calligraphy to Express Emotion from Poems?
by: Liu, Tiancheng, et al.
Published: (2025)
by: Liu, Tiancheng, et al.
Published: (2025)
T2VTree: User-Centered Visual Analytics for Agent-Assisted Thought-to-Video Authoring
by: Zheng, Zhuoyun, et al.
Published: (2026)
by: Zheng, Zhuoyun, et al.
Published: (2026)
Laugh at Your Own Pace: Basic Performance Evaluation of Language Learning Assistance by Adjustment of Video Playback Speeds Based on Laughter Detection
by: Nishida, Naoto, et al.
Published: (2025)
by: Nishida, Naoto, et al.
Published: (2025)
MindCine: Multimodal EEG-to-Video Reconstruction with Large-Scale Pretrained Models
by: Zhou, Tian-Yi, et al.
Published: (2026)
by: Zhou, Tian-Yi, et al.
Published: (2026)
Designing Effective AI Explanations for Misinformation Detection: A Comparative Study of Content, Social, and Combined Explanations
by: Gong, Yeaeun, et al.
Published: (2025)
by: Gong, Yeaeun, et al.
Published: (2025)
Through the Looking-Glass: AI-Mediated Video Communication Reduces Interpersonal Trust and Confidence in Judgments
by: Fernández, Nelson Navajas, et al.
Published: (2026)
by: Fernández, Nelson Navajas, et al.
Published: (2026)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Photoshop Batch Rendering Using Actions for Stylistic Video Editing
by: De La Fuente, Tessa
Published: (2025)
by: De La Fuente, Tessa
Published: (2025)
Chat with AI: The Surprising Turn of Real-time Video Communication from Human to AI
by: Wu, Jiangkai, et al.
Published: (2025)
by: Wu, Jiangkai, et al.
Published: (2025)
Resolution deficits drive simulator sickness and compromise reading performance in virtual environments
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
by: Wang, Bokang, et al.
Published: (2026)
by: Wang, Bokang, et al.
Published: (2026)
The perceptual gap between video see-through displays and natural human vision
by: Wang, Jialin, et al.
Published: (2026)
by: Wang, Jialin, et al.
Published: (2026)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
CvhSlicer 2.0: Immersive and Interactive Visualization of Chinese Visible Human Data in XR Environments
by: Qiu, Yue, et al.
Published: (2025)
by: Qiu, Yue, et al.
Published: (2025)
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration
by: Selvamani, Shaja Arul, et al.
Published: (2025)
by: Selvamani, Shaja Arul, et al.
Published: (2025)
Vidmento: Creating Video Stories Through Context-Aware Expansion With Generative Video
by: Yeh, Catherine, et al.
Published: (2026)
by: Yeh, Catherine, et al.
Published: (2026)
SVFAP: Self-supervised Video Facial Affect Perceiver
by: Sun, Licai, et al.
Published: (2023)
by: Sun, Licai, et al.
Published: (2023)
DiffMesh: A Motion-aware Diffusion Framework for Human Mesh Recovery from Videos
by: Zheng, Ce, et al.
Published: (2023)
by: Zheng, Ce, et al.
Published: (2023)
AffectMachine-Pop: A controllable expert system for real-time pop music generation
by: Agres, Kat R., et al.
Published: (2025)
by: Agres, Kat R., et al.
Published: (2025)
Winds Through Time: Interactive Data Visualization and Physicalization for Paleoclimate Communication
by: Hunter, David, et al.
Published: (2025)
by: Hunter, David, et al.
Published: (2025)
Save It for the "Hot" Day: An LLM-Empowered Visual Analytics System for Heat Risk Management
by: Li, Haobo, et al.
Published: (2024)
by: Li, Haobo, et al.
Published: (2024)
MRATTS: An MR-Based Acupoint Therapy Training System with Real-Time Acupoint Detection and Evaluation Standards
by: Liu, Jiacheng, et al.
Published: (2026)
by: Liu, Jiacheng, et al.
Published: (2026)
Multimodal Digital Sensing of Early-Life Laying Hens: A Pilot Study Integrating Thermal, Acoustic, Optical-Flow and Environmental Data
by: Dhaliwal, Yashan, et al.
Published: (2026)
by: Dhaliwal, Yashan, et al.
Published: (2026)
Similar Items
-
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024) -
Whispering Water: Materializing Human-AI Dialogue as Interactive Ripples
by: Wang, Ruipeng, et al.
Published: (2026) -
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025) -
MambaGesture: Enhancing Co-Speech Gesture Generation with Mamba and Disentangled Multi-Modality Fusion
by: Fu, Chencan, et al.
Published: (2024) -
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
by: Song, Jiarun, et al.
Published: (2026)