Saved in:
| Main Authors: | Liu, Zhaocheng, Yu, Zhiwen, Liu, Xiaoqing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.21797 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Guiding Audio Editing with Audio Language Model
by: Lan, Zitong, et al.
Published: (2025)
by: Lan, Zitong, et al.
Published: (2025)
Learning Source Disentanglement in Neural Audio Codec
by: Bie, Xiaoyu, et al.
Published: (2024)
by: Bie, Xiaoyu, et al.
Published: (2024)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
by: Zhang, Kuiyuan, et al.
Published: (2025)
by: Zhang, Kuiyuan, et al.
Published: (2025)
Multimodal Audio-based Disease Prediction with Transformer-based Hierarchical Fusion Network
by: Cai, Jinjin, et al.
Published: (2024)
by: Cai, Jinjin, et al.
Published: (2024)
Sim2Real Transfer for Audio-Visual Navigation with Frequency-Adaptive Acoustic Field Prediction
by: Chen, Changan, et al.
Published: (2024)
by: Chen, Changan, et al.
Published: (2024)
A Survey of Deep Learning Audio Generation Methods
by: Božić, Matej, et al.
Published: (2024)
by: Božić, Matej, et al.
Published: (2024)
$\texttt{AVROBUSTBENCH}$: Benchmarking the Robustness of Audio-Visual Recognition Models at Test-Time
by: Maharana, Sarthak Kumar, et al.
Published: (2025)
by: Maharana, Sarthak Kumar, et al.
Published: (2025)
RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
by: Zhang, Yuwei, et al.
Published: (2024)
by: Zhang, Yuwei, et al.
Published: (2024)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
by: Chen, Chih-Ning, et al.
Published: (2026)
by: Chen, Chih-Ning, et al.
Published: (2026)
Content Adaptive Front End For Audio Classification
by: Verma, Prateek, et al.
Published: (2023)
by: Verma, Prateek, et al.
Published: (2023)
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
by: Fujita, Yoto, et al.
Published: (2024)
by: Fujita, Yoto, et al.
Published: (2024)
Learning Linearity in Audio Consistency Autoencoders via Implicit Regularization
by: Torres, Bernardo, et al.
Published: (2025)
by: Torres, Bernardo, et al.
Published: (2025)
AudioGenX: Explainability on Text-to-Audio Generative Models
by: Kang, Hyunju, et al.
Published: (2025)
by: Kang, Hyunju, et al.
Published: (2025)
Prompt-guided Precise Audio Editing with Diffusion Models
by: Xu, Manjie, et al.
Published: (2024)
by: Xu, Manjie, et al.
Published: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
by: Wu, Linzhi, et al.
Published: (2026)
by: Wu, Linzhi, et al.
Published: (2026)
LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs
by: Mousavi, Pooneh, et al.
Published: (2025)
by: Mousavi, Pooneh, et al.
Published: (2025)
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
by: Tuncay, Ludovic, et al.
Published: (2025)
by: Tuncay, Ludovic, et al.
Published: (2025)
Investigating Design Choices in Joint-Embedding Predictive Architectures for General Audio Representation Learning
by: Riou, Alain, et al.
Published: (2024)
by: Riou, Alain, et al.
Published: (2024)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation
by: Liu, Teng, et al.
Published: (2026)
by: Liu, Teng, et al.
Published: (2026)
HEAR: Holistic Evaluation of Audio Representations
by: Turian, Joseph, et al.
Published: (2022)
by: Turian, Joseph, et al.
Published: (2022)
Gull: A Generative Multifunctional Audio Codec
by: Luo, Yi, et al.
Published: (2024)
by: Luo, Yi, et al.
Published: (2024)
Speech Diarization and ASR with GMM
by: Sharma, Aayush Kumar, et al.
Published: (2023)
by: Sharma, Aayush Kumar, et al.
Published: (2023)
TACNET: Temporal Audio Source Counting Network
by: Ahmadnejad, Amirreza, et al.
Published: (2023)
by: Ahmadnejad, Amirreza, et al.
Published: (2023)
Training-Free Multi-Step Audio Source Separation
by: Zang, Yongyi, et al.
Published: (2025)
by: Zang, Yongyi, et al.
Published: (2025)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
by: Morocutti, Tobias, et al.
Published: (2025)
by: Morocutti, Tobias, et al.
Published: (2025)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Audio-Based Pedestrian Detection in the Presence of Vehicular Noise
by: Kim, Yonghyun, et al.
Published: (2025)
by: Kim, Yonghyun, et al.
Published: (2025)
Exploring and Applying Audio-Based Sentiment Analysis in Music
by: Jhanji, Etash
Published: (2024)
by: Jhanji, Etash
Published: (2024)
Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
by: Shi, Pujin, et al.
Published: (2024)
by: Shi, Pujin, et al.
Published: (2024)
Text-Queried Audio Source Separation via Hierarchical Modeling
by: Yin, Xinlei, et al.
Published: (2025)
by: Yin, Xinlei, et al.
Published: (2025)
Sparse Autoencoders Make Audio Foundation Models more Explainable
by: Mariotte, Théo, et al.
Published: (2025)
by: Mariotte, Théo, et al.
Published: (2025)
Evaluating Fake Music Detection Performance Under Audio Augmentations
by: Sroka, Tomasz, et al.
Published: (2025)
by: Sroka, Tomasz, et al.
Published: (2025)
Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
PoDAR: Power-Disentangled Audio Representation for Generative Modeling
by: Luebs, Alejandro, et al.
Published: (2026)
by: Luebs, Alejandro, et al.
Published: (2026)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
by: Robinson, David, et al.
Published: (2024)
by: Robinson, David, et al.
Published: (2024)
XAttnMark: Learning Robust Audio Watermarking with Cross-Attention
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Similar Items
-
Guiding Audio Editing with Audio Language Model
by: Lan, Zitong, et al.
Published: (2025) -
Learning Source Disentanglement in Neural Audio Codec
by: Bie, Xiaoyu, et al.
Published: (2024) -
Audio Spatially-Guided Fusion for Audio-Visual Navigation
by: Zhou, Xinyu, et al.
Published: (2026) -
Lightweight Joint Audio-Visual Deepfake Detection via Single-Stream Multi-Modal Learning Framework
by: Zhang, Kuiyuan, et al.
Published: (2025) -
Multimodal Audio-based Disease Prediction with Transformer-based Hierarchical Fusion Network
by: Cai, Jinjin, et al.
Published: (2024)