SeeingSounds: Learning Audio-to-Visual Alignment via Text
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Carnemolla, Simone, Pennisi, Matteo, Russo, Chiara, Palazzo, Simone, Giordano, Daniela, Spampinato, Concetto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
von: Moradi, Morteza, et al.
Veröffentlicht: (2024)
PERL: Parameter Efficient Reasoning in CLIP Latent Space
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
UNBOX: Unveiling Black-box visual models with Natural-language
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026)
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
von: Wang, Jiahua, et al.
Veröffentlicht: (2025)
AV-Edit: Multimodal Generative Sound Effect Editing via Audio-Visual Semantic Joint Control
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
von: Guo, Xinyue, et al.
Veröffentlicht: (2025)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
von: Yang, Jianxuan, et al.
Veröffentlicht: (2025)
Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
von: Senocak, Arda, et al.
Veröffentlicht: (2024)
Sound Sparks Motion: Audio and Text Tuning for Video Editing
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
von: Razlighi, AmirHossein Naghi, et al.
Veröffentlicht: (2026)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
von: Rinaldi, Ivan, et al.
Veröffentlicht: (2026)
Learning Self-Supervised Audio-Visual Representations for Sound Recommendations
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
von: Krishnamurthy, Sudha
Veröffentlicht: (2024)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
Sound2Vision: Generating Diverse Visuals from Audio through Cross-Modal Latent Alignment
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
OCCAM: Open-set Causal Concept explAnation and Ontology induction for black-box vision Models
von: Russo, Chiara Maria, et al.
Veröffentlicht: (2026)
von: Russo, Chiara Maria, et al.
Veröffentlicht: (2026)
READ-Net: Clarifying Emotional Ambiguity via Adaptive Feature Recalibration for Audio-Visual Depression Detection
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
von: Chen, Chenglizhao, et al.
Veröffentlicht: (2026)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
von: Pennisi, Matteo, et al.
Veröffentlicht: (2024)
von: Pennisi, Matteo, et al.
Veröffentlicht: (2024)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
von: Xu, Weihan, et al.
Veröffentlicht: (2025)
Seeing and Hearing: Open-domain Visual-Audio Generation with Diffusion Latent Aligners
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
von: Xing, Yazhou, et al.
Veröffentlicht: (2024)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
von: Pian, Weiguo, et al.
Veröffentlicht: (2026)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
von: Araujo, Edson, et al.
Veröffentlicht: (2026)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
AlignVSR: Audio-Visual Cross-Modal Alignment for Visual Speech Recognition
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
von: Liu, Zehua, et al.
Veröffentlicht: (2024)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
IsoSignVid2Aud: Sign Language Video to Audio Conversion without Text Intermediaries
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
von: Kavediya, Harsh, et al.
Veröffentlicht: (2025)
SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic Binding
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
von: Chen, Mingfei, et al.
Veröffentlicht: (2025)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
von: Kang, Minjae, et al.
Veröffentlicht: (2025)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
von: Araujo, Edson, et al.
Veröffentlicht: (2025)
Continual Audio-Visual Sound Separation
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
von: Pian, Weiguo, et al.
Veröffentlicht: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
von: Liu, Tengfei, et al.
Veröffentlicht: (2026)
PAVAS: Physics-Aware Video-to-Audio Synthesis
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
von: Hyun-Bin, Oh, et al.
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
von: Saleh, Mohamed, et al.
Veröffentlicht: (2026)
Video-to-Audio Generation with Hidden Alignment
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
von: Xu, Manjie, et al.
Veröffentlicht: (2024)
AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache Optimization
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
von: Jiang, Zhonghua, et al.
Veröffentlicht: (2025)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
von: Joo, Seohyun, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Transformer-based Video Saliency Prediction with High Temporal Dimension Decoding
von: Moradi, Morteza, et al.
Veröffentlicht: (2024) -
PERL: Parameter Efficient Reasoning in CLIP Latent Space
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026) -
UNBOX: Unveiling Black-box visual models with Natural-language
von: Carnemolla, Simone, et al.
Veröffentlicht: (2026) -
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
von: Carnemolla, Simone, et al.
Veröffentlicht: (2025) -
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)