AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
Fuente:
arXiv
Saved in:
| Main Authors: | Lau, Kin Wai, Rehman, Yasar Abbas Ur, Po, Lai-Man |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025)
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
by: Lau, Kin Wai, et al.
Published: (2026)
by: Lau, Kin Wai, et al.
Published: (2026)
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
by: Zhou, Wangzixi, et al.
Published: (2026)
by: Zhou, Wangzixi, et al.
Published: (2026)
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
by: Abreu, Wallace, et al.
Published: (2024)
by: Abreu, Wallace, et al.
Published: (2024)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
by: Heo, Hyun-Jun, et al.
Published: (2023)
by: Heo, Hyun-Jun, et al.
Published: (2023)
AudioMorphix: Training-free audio editing with diffusion probabilistic models
by: Liang, Jinhua, et al.
Published: (2025)
by: Liang, Jinhua, et al.
Published: (2025)
Versatile audio-visual learning for emotion recognition
by: Goncalves, Lucas, et al.
Published: (2023)
by: Goncalves, Lucas, et al.
Published: (2023)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
Late fusion ensembles for speech recognition on diverse input audio representations
by: Jezidžić, Marin, et al.
Published: (2024)
by: Jezidžić, Marin, et al.
Published: (2024)
Speech enhancement deep-learning architecture for efficient edge processing
by: Pal, Monisankha, et al.
Published: (2024)
by: Pal, Monisankha, et al.
Published: (2024)
audio2chart: End to End Audio Transcription into playable Guitar Hero charts
by: Tripodi, Riccardo
Published: (2025)
by: Tripodi, Riccardo
Published: (2025)
Online incremental learning for audio classification using a pretrained audio model
by: Mulimani, Manjunath, et al.
Published: (2025)
by: Mulimani, Manjunath, et al.
Published: (2025)
Audio Dialogues: Dialogues dataset for audio and music understanding
by: Goel, Arushi, et al.
Published: (2024)
by: Goel, Arushi, et al.
Published: (2024)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
by: Yang, Yiqing, et al.
Published: (2025)
by: Yang, Yiqing, et al.
Published: (2025)
Scaling up masked audio encoder learning for general audio classification
by: Dinkel, Heinrich, et al.
Published: (2024)
by: Dinkel, Heinrich, et al.
Published: (2024)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
by: Yadav, Sarthak, et al.
Published: (2025)
by: Yadav, Sarthak, et al.
Published: (2025)
Pitch Estimation With Mean Averaging Smoothed Product Spectrum And Musical Consonance Evaluation Using MASP
by: Baskin, Murat Yasar
Published: (2025)
by: Baskin, Murat Yasar
Published: (2025)
WavLM model ensemble for audio deepfake detection
by: Combei, David, et al.
Published: (2024)
by: Combei, David, et al.
Published: (2024)
Cryfish: On deep audio analysis with Large Language Models
by: Mitrofanov, Anton, et al.
Published: (2025)
by: Mitrofanov, Anton, et al.
Published: (2025)
Multiple Hankel matrix rank minimization for audio inpainting
by: Záviška, Pavel, et al.
Published: (2023)
by: Záviška, Pavel, et al.
Published: (2023)
DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation
by: Male, Prabash Reddy, et al.
Published: (2025)
by: Male, Prabash Reddy, et al.
Published: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
by: Yadav, Sarthak, et al.
Published: (2025)
by: Yadav, Sarthak, et al.
Published: (2025)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
by: Wang, Ziqian, et al.
Published: (2025)
by: Wang, Ziqian, et al.
Published: (2025)
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
by: Yang, Chih-Kai, et al.
Published: (2026)
by: Yang, Chih-Kai, et al.
Published: (2026)
SCORE: Scaling audio generation using Standardized COmposite REwards
by: Jung, Jaemin, et al.
Published: (2025)
by: Jung, Jaemin, et al.
Published: (2025)
Towards audio language modeling -- an overview
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
by: Sedukhin, Oleg, et al.
Published: (2026)
by: Sedukhin, Oleg, et al.
Published: (2026)
Towards predicting binaural audio quality in listeners with normal and impaired hearing
by: Biberger, Thomas, et al.
Published: (2025)
by: Biberger, Thomas, et al.
Published: (2025)
Sound event detection with audio-text models and heterogeneous temporal annotations
by: Harju, Manu, et al.
Published: (2025)
by: Harju, Manu, et al.
Published: (2025)
Multi-label audio classification with a noisy zero-shot teacher
by: Braun, Sebastian, et al.
Published: (2024)
by: Braun, Sebastian, et al.
Published: (2024)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
by: Pîrlogeanu, Gabriel, et al.
Published: (2026)
Unmasking real-world audio deepfakes: A data-centric approach
by: Combei, David, et al.
Published: (2025)
by: Combei, David, et al.
Published: (2025)
CardioLive: Empowering Video Streaming with Online Cardiac Monitoring
by: Lyu, Sheng, et al.
Published: (2025)
by: Lyu, Sheng, et al.
Published: (2025)
Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification
by: Gan, Chong-Xin, et al.
Published: (2023)
by: Gan, Chong-Xin, et al.
Published: (2023)
Are audio DeepFake detection models polyglots?
by: Marek, Bartłomiej, et al.
Published: (2024)
by: Marek, Bartłomiej, et al.
Published: (2024)
LDCodec: A high quality neural audio codec with low-complexity decoder
by: Jiang, Jiawei, et al.
Published: (2025)
by: Jiang, Jiawei, et al.
Published: (2025)
Positive and negative sampling strategies for self-supervised learning on audio-video data
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment
by: Zhou, Hao, et al.
Published: (2025)
by: Zhou, Hao, et al.
Published: (2025)
Similar Items
-
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024) -
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2025) -
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
by: Lau, Kin Wai, et al.
Published: (2026) -
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
by: Zhou, Wangzixi, et al.
Published: (2026) -
AEROMamba: An efficient architecture for audio super-resolution using generative adversarial networks and state space models
by: Abreu, Wallace, et al.
Published: (2024)