AxLSTMs: learning self-supervised audio representations with xLSTMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yadav, Sarthak, Theodoridis, Sergios, Tan, Zheng-Hua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2023)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2023)
Exploring bat song syllable representations in self-supervised audio encoders
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
von: Pascu, Octavian, et al.
Veröffentlicht: (2023)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
EDTC: enhance depth of text comprehension in automated audio captioning
von: Tan, Liwen, et al.
Veröffentlicht: (2024)
von: Tan, Liwen, et al.
Veröffentlicht: (2024)
Deep learning based spatial aliasing reduction in beamforming for audio capture
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
von: Guzik, Mateusz, et al.
Veröffentlicht: (2025)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Equivariance-based self-supervised learning for audio signal recovery from clipped measurements
von: Sechaud, Victor, et al.
Veröffentlicht: (2024)
von: Sechaud, Victor, et al.
Veröffentlicht: (2024)
Transformation of audio embeddings into interpretable, concept-based representations
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
von: Chen, Szu-Jui, et al.
Veröffentlicht: (2026)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Tweaking autoregressive methods for inpainting of gaps in audio signals
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
MBCodec:Thorough disentangle for high-fidelity audio compression
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
Real-time implementation of vibrato transfer as an audio effect
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
von: Hyrkas, Jeremy
Veröffentlicht: (2025)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
von: Kühne, Nikolai Lund, et al.
Veröffentlicht: (2025)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
Automated data curation for self-supervised learning in underwater acoustic analysis
von: Hummel, Hilde I, et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I, et al.
Veröffentlicht: (2025)
Regularized autoregressive modeling and its application to audio signal reconstruction
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
FxSearcher: gradient-free text-driven audio transformation
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
von: Ki, Hojoon, et al.
Veröffentlicht: (2025)
How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
STASE: A spatialized text-to-audio synthesis engine for music generation
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
Enhancement by postfiltering for speech and audio coding in ad-hoc sensor networks
von: Das, Sneha, et al.
Veröffentlicht: (2020)
von: Das, Sneha, et al.
Veröffentlicht: (2020)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
GRAM: Spatial general-purpose audio representation models for real-world applications
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Real-Time Scream Detection and Position Estimation for Worker Safety in Construction Sites
von: Gautam, Bikalpa, et al.
Veröffentlicht: (2024)
von: Gautam, Bikalpa, et al.
Veröffentlicht: (2024)
Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025) -
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025) -
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026) -
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024) -
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)