Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Ishikawa, Yuchi, Nakada, Shota, Munakata, Hokuto, Saito, Kazuhiro, Komatsu, Tatsuya, Aoki, Yoshimitsu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ProLAP: Probabilistic Language-Audio Pre-Training
by: Manabe, Toranosuke, et al.
Published: (2025)
by: Manabe, Toranosuke, et al.
Published: (2025)
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024)
by: Munakata, Hokuto, et al.
Published: (2024)
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025)
by: Ishikawa, Yuchi, et al.
Published: (2025)
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
by: Nakada, Shota, et al.
Published: (2024)
by: Nakada, Shota, et al.
Published: (2024)
Pre-training with Synthetic Patterns for Audio
by: Ishikawa, Yuchi, et al.
Published: (2024)
by: Ishikawa, Yuchi, et al.
Published: (2024)
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
by: Munakata, Hokuto, et al.
Published: (2025)
by: Munakata, Hokuto, et al.
Published: (2025)
Audio Fingerprinting with Holographic Reduced Representations
by: Fujita, Yusuke, et al.
Published: (2024)
by: Fujita, Yusuke, et al.
Published: (2024)
Audio-Visual Speech Enhancement for Spatial Audio - Spatial-VisualVoice and the MAVE Database
by: Yaffe, Danielle, et al.
Published: (2025)
by: Yaffe, Danielle, et al.
Published: (2025)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
by: Comunità, Marco, et al.
Published: (2024)
by: Comunità, Marco, et al.
Published: (2024)
On the Audio Hallucinations in Large Audio-Video Language Models
by: Nishimura, Taichi, et al.
Published: (2024)
by: Nishimura, Taichi, et al.
Published: (2024)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
by: Kim, Jongsuk, et al.
Published: (2025)
by: Kim, Jongsuk, et al.
Published: (2025)
Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework
by: Yang, Hsiang-Cheng, et al.
Published: (2026)
by: Yang, Hsiang-Cheng, et al.
Published: (2026)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Cacophony: An Improved Contrastive Audio-Text Model
by: Zhu, Ge, et al.
Published: (2024)
by: Zhu, Ge, et al.
Published: (2024)
AudioEval: Automatic Dual-Perspective and Multi-Dimensional Evaluation of Text-to-Audio-Generation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
by: Papadopoulos, Aristeidis, et al.
Published: (2025)
by: Papadopoulos, Aristeidis, et al.
Published: (2025)
Uncovering the Visual Contribution in Audio-Visual Speech Recognition
by: Lin, Zhaofeng, et al.
Published: (2024)
by: Lin, Zhaofeng, et al.
Published: (2024)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
by: Yi, Mingjing, et al.
Published: (2024)
by: Yi, Mingjing, et al.
Published: (2024)
Multi-View Based Audio Visual Target Speaker Extraction
by: Yang, Peijun, et al.
Published: (2026)
by: Yang, Peijun, et al.
Published: (2026)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
by: Araujo, Edson, et al.
Published: (2025)
by: Araujo, Edson, et al.
Published: (2025)
UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
by: Cao, Yuqin, et al.
Published: (2024)
by: Cao, Yuqin, et al.
Published: (2024)
Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
by: Munakata, Hokuto, et al.
Published: (2024)
by: Munakata, Hokuto, et al.
Published: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
by: Pan, Zexu, et al.
Published: (2025)
by: Pan, Zexu, et al.
Published: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
by: Li, Xiaolou, et al.
Published: (2024)
by: Li, Xiaolou, et al.
Published: (2024)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
by: Jia, Yuhang, et al.
Published: (2025)
by: Jia, Yuhang, et al.
Published: (2025)
Audio-Visual Feature Synchronization for Robust Speech Enhancement in Hearing Aids
by: Saleem, Nasir, et al.
Published: (2025)
by: Saleem, Nasir, et al.
Published: (2025)
Text-based Audio Retrieval by Learning from Similarities between Audio Captions
by: Xie, Huang, et al.
Published: (2024)
by: Xie, Huang, et al.
Published: (2024)
AudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
by: Bai, Jisheng, et al.
Published: (2023)
by: Bai, Jisheng, et al.
Published: (2023)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
by: Zhao, Junqi, et al.
Published: (2024)
by: Zhao, Junqi, et al.
Published: (2024)
Audio-Visual Approach For Multimodal Concurrent Speaker Detection
by: Eliav, Amit, et al.
Published: (2024)
by: Eliav, Amit, et al.
Published: (2024)
Audio Atlas: Visualizing and Exploring Audio Datasets
by: Lanzendörfer, Luca A., et al.
Published: (2024)
by: Lanzendörfer, Luca A., et al.
Published: (2024)
A Generative-First Neural Audio Autoencoder
by: Casebeer, Jonah, et al.
Published: (2026)
by: Casebeer, Jonah, et al.
Published: (2026)
Target Speaker Lipreading by Audio-Visual Self-Distillation Pretraining and Speaker Adaptation
by: Zhang, Jing-Xuan, et al.
Published: (2025)
by: Zhang, Jing-Xuan, et al.
Published: (2025)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
by: Ok, Seaone, et al.
Published: (2026)
by: Ok, Seaone, et al.
Published: (2026)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
by: Chu, Annie, et al.
Published: (2024)
by: Chu, Annie, et al.
Published: (2024)
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
by: Huang, Kuan-Po, et al.
Published: (2025)
by: Huang, Kuan-Po, et al.
Published: (2025)
MOS-FAD: Improving Fake Audio Detection Via Automatic Mean Opinion Score Prediction
by: Zhou, Wangjin, et al.
Published: (2024)
by: Zhou, Wangjin, et al.
Published: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
by: Liu, Huadai, et al.
Published: (2024)
by: Liu, Huadai, et al.
Published: (2024)
A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
by: Fernandes, Jose Geraldo, et al.
Published: (2024)
by: Fernandes, Jose Geraldo, et al.
Published: (2024)
Similar Items
-
ProLAP: Probabilistic Language-Audio Pre-Training
by: Manabe, Toranosuke, et al.
Published: (2025) -
Language-based Audio Moment Retrieval
by: Munakata, Hokuto, et al.
Published: (2024) -
Listening without Looking: Modality Bias in Audio-Visual Captioning
by: Ishikawa, Yuchi, et al.
Published: (2025) -
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
by: Nakada, Shota, et al.
Published: (2024) -
Pre-training with Synthetic Patterns for Audio
by: Ishikawa, Yuchi, et al.
Published: (2024)