AMAuT: A Flexible and Efficient Multiview Audio Transformer Framework Trained from Scratch
Fuente:
arXiv
Saved in:
| Main Authors: | Shao, Weichuang, Liao, Iman Yi, Maul, Tomas Henrique Bode, Chandesa, Tissa |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
by: Shao, Weichuang, et al.
Published: (2025)
by: Shao, Weichuang, et al.
Published: (2025)
An Investigation of Test-time Adaptation for Audio Classification under Background Noise
by: Shao, Weichuang, et al.
Published: (2025)
by: Shao, Weichuang, et al.
Published: (2025)
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
by: Feng, Jiu, et al.
Published: (2024)
by: Feng, Jiu, et al.
Published: (2024)
AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers
by: Yamamoto, Kohei, et al.
Published: (2025)
by: Yamamoto, Kohei, et al.
Published: (2025)
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
by: Wu, Yanru, et al.
Published: (2026)
by: Wu, Yanru, et al.
Published: (2026)
Training-Free Multimodal Guidance for Video to Audio Generation
by: Grassucci, Eleonora, et al.
Published: (2025)
by: Grassucci, Eleonora, et al.
Published: (2025)
Transformer Based Machine Fault Detection From Audio Input
by: Holla, Kiran Voderhobli
Published: (2026)
by: Holla, Kiran Voderhobli
Published: (2026)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
by: Tabassum, Afrina, et al.
Published: (2024)
by: Tabassum, Afrina, et al.
Published: (2024)
Time-Varying Audio Effect Modeling by End-to-End Adversarial Training
by: Bourdin, Yann, et al.
Published: (2025)
by: Bourdin, Yann, et al.
Published: (2025)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
by: Chen, Wenxi, et al.
Published: (2024)
by: Chen, Wenxi, et al.
Published: (2024)
Fast and Flexible Audio Bandwidth Extension via Vocos
by: Sharma, Yatharth
Published: (2026)
by: Sharma, Yatharth
Published: (2026)
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
by: Xiao, Yixuan, et al.
Published: (2026)
by: Xiao, Yixuan, et al.
Published: (2026)
ADNAC: Audio Denoiser using Neural Audio Codec
by: Jimon, Daniel, et al.
Published: (2025)
by: Jimon, Daniel, et al.
Published: (2025)
Towards Trustworthy Audio Deepfake Detection: A Systematic Framework for Diagnosing and Mitigating Gender Bias
by: Fursule, Aishwarya, et al.
Published: (2026)
by: Fursule, Aishwarya, et al.
Published: (2026)
Multiview Canonical Correlation Analysis for Automatic Pathological Speech Detection
by: Kaloga, Yacouba, et al.
Published: (2024)
by: Kaloga, Yacouba, et al.
Published: (2024)
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
by: Wu, Daiqing, et al.
Published: (2026)
by: Wu, Daiqing, et al.
Published: (2026)
PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
by: Vora, Jayneel, et al.
Published: (2024)
by: Vora, Jayneel, et al.
Published: (2024)
SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
by: Ungersböck, Michael, et al.
Published: (2025)
by: Ungersböck, Michael, et al.
Published: (2025)
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
Virtual Consistency for Audio Editing
by: Cervera, Matthieu, et al.
Published: (2025)
by: Cervera, Matthieu, et al.
Published: (2025)
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
by: Yu, Tao, et al.
Published: (2026)
by: Yu, Tao, et al.
Published: (2026)
Improving Audio Classification by Transitioning from Zero- to Few-Shot
by: Taylor, James, et al.
Published: (2025)
by: Taylor, James, et al.
Published: (2025)
Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
PACE: Pretrained Audio Continual Learning
by: Li, Chang, et al.
Published: (2026)
by: Li, Chang, et al.
Published: (2026)
Segmentwise Pruning in Audio-Language Models
by: Gibier, Marcel, et al.
Published: (2025)
by: Gibier, Marcel, et al.
Published: (2025)
Adapting Neural Audio Codecs to EEG
by: Kastrati, Ard, et al.
Published: (2025)
by: Kastrati, Ard, et al.
Published: (2025)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
by: Sinha, Anshuman, et al.
Published: (2024)
by: Sinha, Anshuman, et al.
Published: (2024)
Investigating Modality Contribution in Audio LLMs for Music
by: Morais, Giovana, et al.
Published: (2025)
by: Morais, Giovana, et al.
Published: (2025)
APEX: Audio Prototype EXplanations for Classification Tasks
by: Kawa, Piotr, et al.
Published: (2026)
by: Kawa, Piotr, et al.
Published: (2026)
Audio Super-Resolution with Latent Bridge Models
by: Li, Chang, et al.
Published: (2025)
by: Li, Chang, et al.
Published: (2025)
SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
by: Dellali, Amir, et al.
Published: (2025)
by: Dellali, Amir, et al.
Published: (2025)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
by: Bruford, Fred, et al.
Published: (2024)
by: Bruford, Fred, et al.
Published: (2024)
Semantic-Aware Confidence Calibration for Automated Audio Captioning
by: Dunker, Lucas, et al.
Published: (2025)
by: Dunker, Lucas, et al.
Published: (2025)
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
How to Count Coughs: An Event-Based Framework for Evaluating Automatic Cough Detection Algorithm Performance
by: Orlandic, Lara, et al.
Published: (2024)
by: Orlandic, Lara, et al.
Published: (2024)
A Human-Inspired Decoupled Architecture for Efficient Audio Representation Learning
by: Kawano, Harunori, et al.
Published: (2026)
by: Kawano, Harunori, et al.
Published: (2026)
Audio Transformers
by: Verma, Prateek, et al.
Published: (2021)
by: Verma, Prateek, et al.
Published: (2021)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
FastWave: Optimized Diffusion Model for Audio Super-Resolution
by: Kuznetsov, Nikita, et al.
Published: (2026)
by: Kuznetsov, Nikita, et al.
Published: (2026)
Audio-Visual Continual Test-Time Adaptation without Forgetting
by: Maharana, Sarthak Kumar, et al.
Published: (2026)
by: Maharana, Sarthak Kumar, et al.
Published: (2026)
Similar Items
-
DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation
by: Shao, Weichuang, et al.
Published: (2025) -
An Investigation of Test-time Adaptation for Audio Classification under Background Noise
by: Shao, Weichuang, et al.
Published: (2025) -
From Coarse to Fine: Efficient Training for Audio Spectrogram Transformers
by: Feng, Jiu, et al.
Published: (2024) -
AaSP: Aliasing-aware Self-Supervised Pre-Training for Audio Spectrogram Transformers
by: Yamamoto, Kohei, et al.
Published: (2025) -
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
by: Wu, Yanru, et al.
Published: (2026)