When Vision Speaks for Sound
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Xiaofei, Mo, Wenjie Jacky, Fu, Xingyu, Cai, Rui, Zhu, Tinghui, Li, Wendi, Xie, Yanan, Chen, Muhao, Qi, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
von: Zhu, Boyu, et al.
Veröffentlicht: (2025)
UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
von: Chen, Zihao, et al.
Veröffentlicht: (2024)
von: Chen, Zihao, et al.
Veröffentlicht: (2024)
Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
Learning Adaptive Reasoning Paths for Efficient Visual Reasoning
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
von: Huang, Yixu, et al.
Veröffentlicht: (2026)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
von: Kim, Inho, et al.
Veröffentlicht: (2025)
von: Kim, Inho, et al.
Veröffentlicht: (2025)
ReelWave: Multi-Agentic Movie Sound Generation through Multimodal LLM Conversation
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
High-Quality Sound Separation Across Diverse Categories via Visually-Guided Generative Modeling
von: Huang, Chao, et al.
Veröffentlicht: (2025)
von: Huang, Chao, et al.
Veröffentlicht: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Multi-scale Multi-instance Visual Sound Localization and Segmentation
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
von: Mo, Shentong, et al.
Veröffentlicht: (2024)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
Differentiable Room Acoustic Rendering with Multi-View Vision Priors
von: Jin, Derong, et al.
Veröffentlicht: (2025)
von: Jin, Derong, et al.
Veröffentlicht: (2025)
From Vision to Sound: Advancing Audio Anomaly Detection with Vision-Based Algorithms
von: Barusco, Manuel, et al.
Veröffentlicht: (2025)
von: Barusco, Manuel, et al.
Veröffentlicht: (2025)
Video Models Can Reason with Verifiable Rewards
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2026)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
Gotta Hear Them All: Towards Sound Source Aware Audio Generation
von: Guo, Wei, et al.
Veröffentlicht: (2024)
von: Guo, Wei, et al.
Veröffentlicht: (2024)
UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech Avatars
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2026)
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
Continual Audio-Visual Sound Separation
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
MOVA: Towards Scalable and Synchronized Video-Audio Generation
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals
von: Liang, Xiaoxuan, et al.
Veröffentlicht: (2024)
von: Liang, Xiaoxuan, et al.
Veröffentlicht: (2024)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits
von: Li, Kai, et al.
Veröffentlicht: (2022)
von: Li, Kai, et al.
Veröffentlicht: (2022)
LTA-L2S: Lexical Tone-Aware Lip-to-Speech Synthesis for Mandarin with Cross-Lingual Transfer Learning
von: Yang, Kang, et al.
Veröffentlicht: (2025)
von: Yang, Kang, et al.
Veröffentlicht: (2025)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts
von: Li, You, et al.
Veröffentlicht: (2026)
von: Li, You, et al.
Veröffentlicht: (2026)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
von: Khanal, Subash, et al.
Veröffentlicht: (2025)
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos
von: Chen, Changan, et al.
Veröffentlicht: (2024)
von: Chen, Changan, et al.
Veröffentlicht: (2024)
SemiPL: A Semi-supervised Method for Event Sound Source Localization
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
Hierarchical Codec Diffusion for Video-to-Speech Generation
von: Ye, Jiaxin, et al.
Veröffentlicht: (2026)
von: Ye, Jiaxin, et al.
Veröffentlicht: (2026)
Semantic Noise Reduction via Teacher-Guided Dual-Path Audio-Visual Representation Learning
von: Wang, Linge, et al.
Veröffentlicht: (2026)
von: Wang, Linge, et al.
Veröffentlicht: (2026)
Segmenting Collision Sound Sources in Egocentric Videos
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
von: Parida, Kranti Kumar, et al.
Veröffentlicht: (2025)
Seeing Sound: Assembling Sounds from Visuals for Audio-to-Image Generation
von: Petermann, Darius, et al.
Veröffentlicht: (2025)
von: Petermann, Darius, et al.
Veröffentlicht: (2025)
Multi-View Spectrogram Transformer for Respiratory Sound Classification
von: He, Wentao, et al.
Veröffentlicht: (2023)
von: He, Wentao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
OmniGuard: Unified Omni-Modal Guardrails with Deliberate Reasoning
von: Zhu, Boyu, et al.
Veröffentlicht: (2025) -
UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025) -
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026) -
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
von: Guan, Kaisi, et al.
Veröffentlicht: (2025) -
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
von: Chen, Zihao, et al.
Veröffentlicht: (2024)