Saved in:
| Main Authors: | Ruff, Zachary J., Lesmeister, Damon B. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.14864 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
by: Liu, Che, et al.
Published: (2024)
by: Liu, Che, et al.
Published: (2024)
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
by: Apostolidis, Konstantinos, et al.
Published: (2024)
by: Apostolidis, Konstantinos, et al.
Published: (2024)
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)
by: Kefalas, Triantafyllos, et al.
Published: (2023)
Audio-Visual Talker Localization in Video for Spatial Sound Reproduction
by: Berghi, Davide, et al.
Published: (2024)
by: Berghi, Davide, et al.
Published: (2024)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Automated Bioacoustic Monitoring for South African Bird Species on Unlabeled Data
by: Doell, Michael, et al.
Published: (2024)
by: Doell, Michael, et al.
Published: (2024)
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
A-JEPA: Joint-Embedding Predictive Architecture Can Listen
by: Fei, Zhengcong, et al.
Published: (2023)
by: Fei, Zhengcong, et al.
Published: (2023)
Learning to Highlight Audio by Watching Movies
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
SoundCam: A Dataset for Finding Humans Using Room Acoustics
by: Wang, Mason, et al.
Published: (2023)
by: Wang, Mason, et al.
Published: (2023)
Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
by: Li, Zeyu, et al.
Published: (2024)
by: Li, Zeyu, et al.
Published: (2024)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
by: Park, Sooyoung, et al.
Published: (2025)
by: Park, Sooyoung, et al.
Published: (2025)
A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
by: Juanola, Xavier, et al.
Published: (2024)
by: Juanola, Xavier, et al.
Published: (2024)
Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
by: Shou, Yuntao, et al.
Published: (2023)
by: Shou, Yuntao, et al.
Published: (2023)
Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization
by: Zhang, Haomin, et al.
Published: (2025)
by: Zhang, Haomin, et al.
Published: (2025)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
by: Ryu, Hyeonggon, et al.
Published: (2025)
by: Ryu, Hyeonggon, et al.
Published: (2025)
Enhancing Dance-to-Music Generation via Negative Conditioning Latent Diffusion Model
by: Sun, Changchang, et al.
Published: (2025)
by: Sun, Changchang, et al.
Published: (2025)
SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
by: Park, Young-Hu, et al.
Published: (2025)
by: Park, Young-Hu, et al.
Published: (2025)
UWAV: Uncertainty-weighted Weakly-supervised Audio-Visual Video Parsing
by: Lai, Yung-Hsuan, et al.
Published: (2025)
by: Lai, Yung-Hsuan, et al.
Published: (2025)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
CCStereo: Audio-Visual Contextual and Contrastive Learning for Binaural Audio Generation
by: Chen, Yuanhong, et al.
Published: (2025)
by: Chen, Yuanhong, et al.
Published: (2025)
From Detection to Correction: Backdoor-Resilient Face Recognition via Vision-Language Trigger Detection and Noise-Based Neutralization
by: Wahida, Farah, et al.
Published: (2025)
by: Wahida, Farah, et al.
Published: (2025)
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
by: Kwon, Mingi, et al.
Published: (2025)
by: Kwon, Mingi, et al.
Published: (2025)
ZeroSep: Separate Anything in Audio with Zero Training
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
As Good as It KAN Get: High-Fidelity Audio Representation
by: Marszałek, Patryk, et al.
Published: (2025)
by: Marszałek, Patryk, et al.
Published: (2025)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
NaturalL2S: End-to-End High-quality Multispeaker Lip-to-Speech Synthesis with Differential Digital Signal Processing
by: Liang, Yifan, et al.
Published: (2025)
by: Liang, Yifan, et al.
Published: (2025)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
by: Zhang, Zhedong, et al.
Published: (2025)
by: Zhang, Zhedong, et al.
Published: (2025)
Improving Acoustic Scene Classification with City Features
by: Cai, Yiqiang, et al.
Published: (2025)
by: Cai, Yiqiang, et al.
Published: (2025)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
by: Fu, Chaoyou, et al.
Published: (2025)
by: Fu, Chaoyou, et al.
Published: (2025)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
by: Cappellazzo, Umberto, et al.
Published: (2025)
by: Cappellazzo, Umberto, et al.
Published: (2025)
Voice Pathology Detection Using Phonation
by: Siva, Sri Raksha, et al.
Published: (2025)
by: Siva, Sri Raksha, et al.
Published: (2025)
Dynamic Derivation and Elimination: Audio Visual Segmentation with Enhanced Audio Semantics
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
How Would It Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
by: Saad, Mahnoor Fatima, et al.
Published: (2025)
by: Saad, Mahnoor Fatima, et al.
Published: (2025)
Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts
by: Farhadipour, Aref, et al.
Published: (2025)
by: Farhadipour, Aref, et al.
Published: (2025)
Similar Items
-
Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
by: Liu, Che, et al.
Published: (2024) -
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024) -
Visual and audio scene classification for detecting discrepancies in video: a baseline method and experimental protocol
by: Apostolidis, Konstantinos, et al.
Published: (2024) -
Audio-visual video-to-speech synthesis with synthesized input audio
by: Kefalas, Triantafyllos, et al.
Published: (2023) -
Large-scale unsupervised audio pre-training for video-to-speech synthesis
by: Kefalas, Triantafyllos, et al.
Published: (2023)