ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Junyu, Wang, Tianrui, Ge, Meng, Wang, Longbiao, Dang, Jianwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
von: Wang, Junyu, et al.
Veröffentlicht: (2024)
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
von: Wang, Junyu, et al.
Veröffentlicht: (2025)
InstructAudio: Unified speech and music generation with natural language instruction
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2024)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
Enriching Multimodal Sentiment Analysis through Textual Emotional Descriptions of Visual-Audio Content
von: Wu, Sheng, et al.
Veröffentlicht: (2024)
von: Wu, Sheng, et al.
Veröffentlicht: (2024)
UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
von: Shu, Yuchun, et al.
Veröffentlicht: (2024)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
von: Cui, Zhongjian, et al.
Veröffentlicht: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
ZMM-TTS: Zero-shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-supervised Discrete Speech Representations
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
von: Gong, Cheng, et al.
Veröffentlicht: (2023)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
von: Luong, Justin, et al.
Veröffentlicht: (2025)
von: Luong, Justin, et al.
Veröffentlicht: (2025)
Robust Audio Anti-Spoofing with Fusion-Reconstruction Learning on Multi-Order Spectrograms
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
von: Wen, Penghui, et al.
Veröffentlicht: (2023)
Audio Contrastive-based Fine-tuning: Decoupling Representation Learning and Classification
von: Wang, Yang, et al.
Veröffentlicht: (2023)
von: Wang, Yang, et al.
Veröffentlicht: (2023)
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
von: Feng, Jiu, et al.
Veröffentlicht: (2024)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
von: Li, Yuxin, et al.
Veröffentlicht: (2025)
Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of Deep Learning Models
von: Pham, Lam, et al.
Veröffentlicht: (2024)
von: Pham, Lam, et al.
Veröffentlicht: (2024)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
ASM: Audio Spectrogram Mixer
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
von: Li, Jilong, et al.
Veröffentlicht: (2024)
von: Li, Jilong, et al.
Veröffentlicht: (2024)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
von: Han, Bing, et al.
Veröffentlicht: (2026)
von: Han, Bing, et al.
Veröffentlicht: (2026)
Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
von: Xie, Yuankun, et al.
Veröffentlicht: (2024)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
Efficient Emotion and Speaker Adaptation in LLM-Based TTS via Characteristic-Specific Partial Fine-Tuning
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
von: He, Peize, et al.
Veröffentlicht: (2025)
von: He, Peize, et al.
Veröffentlicht: (2025)
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
von: Zhou, Li, et al.
Veröffentlicht: (2026)
von: Zhou, Li, et al.
Veröffentlicht: (2026)
Enhancing Partially Spoofed Audio Localization with Boundary-aware Attention Mechanism
von: Zhong, Jiafeng, et al.
Veröffentlicht: (2024)
von: Zhong, Jiafeng, et al.
Veröffentlicht: (2024)
Music Genre Classification: A Comparative Analysis of CNN and XGBoost Approaches with Mel-frequency cepstral coefficients and Mel Spectrograms
von: Meng, Yigang
Veröffentlicht: (2024)
von: Meng, Yigang
Veröffentlicht: (2024)
HierCon: Hierarchical Contrastive Attention for Audio Deepfake Detection
von: Liang, Zhili Nicholas, et al.
Veröffentlicht: (2026)
von: Liang, Zhili Nicholas, et al.
Veröffentlicht: (2026)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
In-the-wild Audio Spatialization with Flexible Text-guided Localization
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025)
von: Shao, Nian, et al.
Veröffentlicht: (2025)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
von: Yan, Canxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2024) -
LORT: Locally Refined Convolution and Taylor Transformer for Monaural Speech Enhancement
von: Wang, Junyu, et al.
Veröffentlicht: (2025) -
Pay More Attention To Audio: Mitigating Imbalance of Cross-Modal Attention in Large Audio Language Models
von: Wang, Junyu, et al.
Veröffentlicht: (2025) -
InstructAudio: Unified speech and music generation with natural language instruction
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025) -
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)