GRAM: Spatial general-purpose audio representation models for real-world applications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yuksel, Goksenin, van Gerven, Marcel, van der Heijden, Kiki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GRAM: Spatial general-purpose audio representations for real-world environments
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2026)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2026)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
von: Ledder, Wessel, et al.
Veröffentlicht: (2024)
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
Making deep neural networks work for medical audio: representation, compression and domain adaptation
von: Onu, Charles C
Veröffentlicht: (2025)
von: Onu, Charles C
Veröffentlicht: (2025)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
Mellow: a small audio language model for reasoning
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
von: Ratnarajah, Anton Jeran
Veröffentlicht: (2024)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
ADIFF: Explaining audio difference using natural language
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Exploring bat song syllable representations in self-supervised audio encoders
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Discriminating real and synthetic super-resolved audio samples using embedding-based classifiers
von: Silaev, Mikhail, et al.
Veröffentlicht: (2026)
von: Silaev, Mikhail, et al.
Veröffentlicht: (2026)
Sustaining model performance for covid-19 detection from dynamic audio data: Development and evaluation of a comprehensive drift-adaptive framework
von: Ganitidis, Theofanis, et al.
Veröffentlicht: (2024)
von: Ganitidis, Theofanis, et al.
Veröffentlicht: (2024)
AISTAT lab system for DCASE2025 Task6: Language-based audio retrieval
von: Kim, Hyun Jun, et al.
Veröffentlicht: (2025)
von: Kim, Hyun Jun, et al.
Veröffentlicht: (2025)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Automated data curation for self-supervised learning in underwater acoustic analysis
von: Hummel, Hilde I, et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I, et al.
Veröffentlicht: (2025)
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
Regularized autoregressive modeling and its application to audio signal reconstruction
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
von: Mokrý, Ondřej, et al.
Veröffentlicht: (2024)
Non-autoregressive real-time Accent Conversion model with voice cloning
von: Nechaev, Vladimir, et al.
Veröffentlicht: (2024)
von: Nechaev, Vladimir, et al.
Veröffentlicht: (2024)
BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
von: Kuang, Sheng, et al.
Veröffentlicht: (2022)
Linear Time Complexity Conformers with SummaryMixing for Streaming Speech Recognition
von: Parcollet, Titouan, et al.
Veröffentlicht: (2024)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2024)
Real-world Music Plagiarism Detection With Music Segment Transcription System
von: Go, Seonghyeon
Veröffentlicht: (2025)
von: Go, Seonghyeon
Veröffentlicht: (2025)
AxLSTMs: learning self-supervised audio representations with xLSTMs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Recomposer: Event-roll-guided generative audio editing
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
ViSAGe: Video-to-Spatial Audio Generation
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
SpatialEmb: Extract and Encode Spatial Information for 1-Stage Multi-channel Multi-speaker ASR on Arbitrary Microphone Arrays
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
von: Shao, Yiwen, et al.
Veröffentlicht: (2026)
In-the-wild Audio Spatialization with Flexible Text-guided Localization
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
von: Wu, Shaohang, et al.
Veröffentlicht: (2026)
Forensic deepfake audio detection using segmental speech features
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GRAM: Spatial general-purpose audio representations for real-world environments
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2026) -
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025) -
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025) -
Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
von: Ledder, Wessel, et al.
Veröffentlicht: (2024) -
Self-supervised learning method using multiple sampling strategies for general-purpose audio representation
von: Kuroyanagi, Ibuki, et al.
Veröffentlicht: (2025)