A-JEPA: Joint-Embedding Predictive Architecture Can Listen
Fuente:
arXiv
Saved in:
| Main Authors: | Fei, Zhengcong, Fan, Mingyuan, Huang, Junshi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLUX that Plays Music
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
Music Consistency Models
by: Fei, Zhengcong, et al.
Published: (2024)
by: Fei, Zhengcong, et al.
Published: (2024)
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
by: Liu, Xi, et al.
Published: (2024)
by: Liu, Xi, et al.
Published: (2024)
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
by: Korbar, Bruno, et al.
Published: (2024)
by: Korbar, Bruno, et al.
Published: (2024)
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
by: Huang, Jinhe, et al.
Published: (2025)
by: Huang, Jinhe, et al.
Published: (2025)
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
Aligned Better, Listen Better for Audio-Visual Large Language Models
by: Guo, Yuxin, et al.
Published: (2025)
by: Guo, Yuxin, et al.
Published: (2025)
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching
by: Kwon, Mingi, et al.
Published: (2025)
by: Kwon, Mingi, et al.
Published: (2025)
Audio-Visual Person Verification based on Recursive Fusion of Joint Cross-Attention
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Recursive Joint Cross-Modal Attention for Multimodal Fusion in Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
UniSync: A Unified Framework for Audio-Visual Synchronization
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
by: Khanal, Subash, et al.
Published: (2024)
by: Khanal, Subash, et al.
Published: (2024)
Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
by: Yang, Shu-wen, et al.
Published: (2025)
by: Yang, Shu-wen, et al.
Published: (2025)
DualTalk: Dual-Speaker Interaction for 3D Talking Head Conversations
by: Peng, Ziqiao, et al.
Published: (2025)
by: Peng, Ziqiao, et al.
Published: (2025)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
by: Kushwaha, Saksham Singh, et al.
Published: (2024)
by: Kushwaha, Saksham Singh, et al.
Published: (2024)
SoundWeaver: Semantic Warm-Starting for Text-to-Audio Diffusion Serving
by: Barik, Ayush, et al.
Published: (2026)
by: Barik, Ayush, et al.
Published: (2026)
ZeroSep: Separate Anything in Audio with Zero Training
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
by: Du, Jiarong, et al.
Published: (2025)
by: Du, Jiarong, et al.
Published: (2025)
Modeling and Driving Human Body Soundfields through Acoustic Primitives
by: Huang, Chao, et al.
Published: (2024)
by: Huang, Chao, et al.
Published: (2024)
High-Quality Visually-Guided Sound Separation from Diverse Categories
by: Huang, Chao, et al.
Published: (2023)
by: Huang, Chao, et al.
Published: (2023)
AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition
by: Xue, Junxiao, et al.
Published: (2025)
by: Xue, Junxiao, et al.
Published: (2025)
ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
by: Cheng, Yongkang, et al.
Published: (2024)
by: Cheng, Yongkang, et al.
Published: (2024)
Separate to Collaborate: Dual-Stream Diffusion Model for Coordinated Piano Hand Motion Synthesis
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
Learning to Highlight Audio by Watching Movies
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
by: Fu, Chaoyou, et al.
Published: (2025)
by: Fu, Chaoyou, et al.
Published: (2025)
Straight Through Gumbel Softmax Estimator based Bimodal Neural Architecture Search for Audio-Visual Deepfake Detection
by: PN, Aravinda Reddy, et al.
Published: (2024)
by: PN, Aravinda Reddy, et al.
Published: (2024)
OmniAudio: Generating Spatial Audio from 360-Degree Video
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
SoundCam: A Dataset for Finding Humans Using Room Acoustics
by: Wang, Mason, et al.
Published: (2023)
by: Wang, Mason, et al.
Published: (2023)
pycnet-audio: A Python package to support bioacoustics data processing
by: Ruff, Zachary J., et al.
Published: (2025)
by: Ruff, Zachary J., et al.
Published: (2025)
Oceanship: A Large-Scale Dataset for Underwater Audio Target Recognition
by: Li, Zeyu, et al.
Published: (2024)
by: Li, Zeyu, et al.
Published: (2024)
A Critical Assessment of Visual Sound Source Localization Models Including Negative Audio
by: Juanola, Xavier, et al.
Published: (2024)
by: Juanola, Xavier, et al.
Published: (2024)
Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization
by: Park, Sooyoung, et al.
Published: (2025)
by: Park, Sooyoung, et al.
Published: (2025)
Towards Reliable Audio Deepfake Attribution and Model Recognition: A Multi-Level Autoencoder-Based Framework
by: Di Pierno, Andrea, et al.
Published: (2025)
by: Di Pierno, Andrea, et al.
Published: (2025)
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
A Low-rank Matching Attention based Cross-modal Feature Fusion Method for Conversational Emotion Recognition
by: Shou, Yuntao, et al.
Published: (2023)
by: Shou, Yuntao, et al.
Published: (2023)
Enhance Generation Quality of Flow Matching V2A Model via Multi-Step CoT-Like Guidance and Combined Preference Optimization
by: Zhang, Haomin, et al.
Published: (2025)
by: Zhang, Haomin, et al.
Published: (2025)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
by: Li, Cancan, et al.
Published: (2025)
by: Li, Cancan, et al.
Published: (2025)
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024)
by: Waligora, Paul, et al.
Published: (2024)
Similar Items
-
FLUX that Plays Music
by: Fei, Zhengcong, et al.
Published: (2024) -
Music Consistency Models
by: Fei, Zhengcong, et al.
Published: (2024) -
CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation
by: Liu, Xi, et al.
Published: (2024) -
Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
by: Korbar, Bruno, et al.
Published: (2024) -
Inter-Diffusion Generation Model of Speakers and Listeners for Effective Communication
by: Huang, Jinhe, et al.
Published: (2025)