Representation-Regularized Convolutional Audio Transformer for Audio Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Han, Bing, Zhou, Chushu, Yang, Yifan, Wang, Wei, Li, Chenda, Zhang, Wangyou, Qian, Yanmin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
di: Zhang, Leying, et al.
Pubblicazione: (2024)
di: Zhang, Leying, et al.
Pubblicazione: (2024)
Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
di: Zheng, Xinhu, et al.
Pubblicazione: (2024)
di: Zheng, Xinhu, et al.
Pubblicazione: (2024)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
di: Zhang, Leying, et al.
Pubblicazione: (2026)
di: Zhang, Leying, et al.
Pubblicazione: (2026)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
di: Chen, Liyang, et al.
Pubblicazione: (2026)
di: Chen, Liyang, et al.
Pubblicazione: (2026)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
di: Zhou, Xinyu, et al.
Pubblicazione: (2026)
Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models
di: Yang, Wanqi, et al.
Pubblicazione: (2024)
di: Yang, Wanqi, et al.
Pubblicazione: (2024)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
di: Song, Zirui, et al.
Pubblicazione: (2025)
di: Song, Zirui, et al.
Pubblicazione: (2025)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
di: Luong, Justin, et al.
Pubblicazione: (2025)
di: Luong, Justin, et al.
Pubblicazione: (2025)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
di: Han, Bing, et al.
Pubblicazione: (2025)
di: Han, Bing, et al.
Pubblicazione: (2025)
LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
di: Zhang, Zhisheng, et al.
Pubblicazione: (2026)
di: Zhang, Zhisheng, et al.
Pubblicazione: (2026)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
di: Zhao, Junqi, et al.
Pubblicazione: (2025)
di: Zhao, Junqi, et al.
Pubblicazione: (2025)
Cross-Domain Audio Deepfake Detection: Dataset and Analysis
di: Li, Yuang, et al.
Pubblicazione: (2024)
di: Li, Yuang, et al.
Pubblicazione: (2024)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
di: Lin, Jiaju, et al.
Pubblicazione: (2024)
di: Lin, Jiaju, et al.
Pubblicazione: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
AAT: Adapting Audio Transformer for Various Acoustics Recognition Tasks
di: Liang, Yun, et al.
Pubblicazione: (2024)
di: Liang, Yun, et al.
Pubblicazione: (2024)
Harder or Different? Understanding Generalization of Audio Deepfake Detection
di: Müller, Nicolas M., et al.
Pubblicazione: (2024)
di: Müller, Nicolas M., et al.
Pubblicazione: (2024)
Raw Audio Classification with Cosine Convolutional Neural Network (CosCovNN)
di: Haque, Kazi Nazmul, et al.
Pubblicazione: (2024)
di: Haque, Kazi Nazmul, et al.
Pubblicazione: (2024)
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
di: He, Peize, et al.
Pubblicazione: (2025)
di: He, Peize, et al.
Pubblicazione: (2025)
Audio Atlas: Visualizing and Exploring Audio Datasets
di: Lanzendörfer, Luca A., et al.
Pubblicazione: (2024)
di: Lanzendörfer, Luca A., et al.
Pubblicazione: (2024)
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
di: Yang, Haoyun, et al.
Pubblicazione: (2026)
di: Yang, Haoyun, et al.
Pubblicazione: (2026)
UltraEval-Audio: A Unified Framework for Comprehensive Evaluation of Audio Foundation Models
di: Shi, Qundong, et al.
Pubblicazione: (2026)
di: Shi, Qundong, et al.
Pubblicazione: (2026)
Towards Controllable Audio Texture Morphing
di: Gupta, Chitralekha, et al.
Pubblicazione: (2023)
di: Gupta, Chitralekha, et al.
Pubblicazione: (2023)
Generalizable Audio Spoofing Detection using Non-Semantic Representations
di: Das, Arnab, et al.
Pubblicazione: (2025)
di: Das, Arnab, et al.
Pubblicazione: (2025)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
di: Li, Chenda, et al.
Pubblicazione: (2024)
di: Li, Chenda, et al.
Pubblicazione: (2024)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
di: Singh, Arshdeep, et al.
Pubblicazione: (2025)
di: Singh, Arshdeep, et al.
Pubblicazione: (2025)
SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition
di: Wu, Yihan, et al.
Pubblicazione: (2024)
di: Wu, Yihan, et al.
Pubblicazione: (2024)
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
di: Yuan, Yi, et al.
Pubblicazione: (2025)
di: Yuan, Yi, et al.
Pubblicazione: (2025)
All That Glitters Is Not Audio: Rethinking Text Priors and Audio Reliance in Audio-Language Evaluation
di: Foo, Leonardo Haw-Yang, et al.
Pubblicazione: (2026)
di: Foo, Leonardo Haw-Yang, et al.
Pubblicazione: (2026)
ElasticAST: An Audio Spectrogram Transformer for All Length and Resolutions
di: Feng, Jiu, et al.
Pubblicazione: (2024)
di: Feng, Jiu, et al.
Pubblicazione: (2024)
HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
di: Hu, Xuyi, et al.
Pubblicazione: (2025)
di: Hu, Xuyi, et al.
Pubblicazione: (2025)
SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
di: Mei, Xinhao, et al.
Pubblicazione: (2026)
di: Mei, Xinhao, et al.
Pubblicazione: (2026)
Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks
di: Atif, Youness
Pubblicazione: (2025)
di: Atif, Youness
Pubblicazione: (2025)
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification
di: Liu, Bei, et al.
Pubblicazione: (2024)
di: Liu, Bei, et al.
Pubblicazione: (2024)
Improving Design of Input Condition Invariant Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
di: Zhang, Wangyou, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling
di: Zhang, Leying, et al.
Pubblicazione: (2024) -
Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
di: Zheng, Xinhu, et al.
Pubblicazione: (2024) -
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
di: Wang, Wei, et al.
Pubblicazione: (2025) -
Beyond Performance Plateaus: A Comprehensive Study on Scalability in Speech Enhancement
di: Zhang, Wangyou, et al.
Pubblicazione: (2024) -
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
di: Zhang, Leying, et al.
Pubblicazione: (2026)