Better audio representations are more brain-like: linking model-brain alignment with performance in downstream auditory tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pepino, Leonardo, Riera, Pablo, Kamienkowski, Juan, Ferrer, Luciana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024)
Training chord recognition models on artificially generated audio
von: Majchrzak, Martyna, et al.
Veröffentlicht: (2025)
von: Majchrzak, Martyna, et al.
Veröffentlicht: (2025)
Transformation of audio embeddings into interpretable, concept-based representations
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
Benchmarking Time-localized Explanations for Audio Classification Models
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
von: Gauder, Lara, et al.
Veröffentlicht: (2024)
von: Gauder, Lara, et al.
Veröffentlicht: (2024)
Investigating self-supervised representations for audio-visual deepfake detection
von: Boldisor, Dragos-Alexandru, et al.
Veröffentlicht: (2025)
von: Boldisor, Dragos-Alexandru, et al.
Veröffentlicht: (2025)
Exploring bat song syllable representations in self-supervised audio encoders
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Enabling automatic transcription of child-centered audio recordings from real-world environments
von: Kocharov, Daniil, et al.
Veröffentlicht: (2025)
von: Kocharov, Daniil, et al.
Veröffentlicht: (2025)
IsoNet: Spatially-aware audio-visual target speech extraction in complex acoustic environments
von: Padhya, Dinanath, et al.
Veröffentlicht: (2026)
von: Padhya, Dinanath, et al.
Veröffentlicht: (2026)
Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
von: Souza, Gabriel, et al.
Veröffentlicht: (2024)
von: Souza, Gabriel, et al.
Veröffentlicht: (2024)
A contrastive-learning approach for auditory attention detection
von: Bajestan, Seyed Ali Alavi, et al.
Veröffentlicht: (2024)
von: Bajestan, Seyed Ali Alavi, et al.
Veröffentlicht: (2024)
The silence of the weights: a structural pruning strategy for attention-based audio signal architectures with second order metrics
von: Diecidue, Andrea, et al.
Veröffentlicht: (2025)
von: Diecidue, Andrea, et al.
Veröffentlicht: (2025)
Towards generalizing deep-audio fake detection networks
von: Gasenzer, Konstantin, et al.
Veröffentlicht: (2023)
von: Gasenzer, Konstantin, et al.
Veröffentlicht: (2023)
Testing chatbots on the creation of encoders for audio conditioned image generation
von: León, Jorge E., et al.
Veröffentlicht: (2025)
von: León, Jorge E., et al.
Veröffentlicht: (2025)
Unsupervised outlier detection to improve bird audio dataset labels
von: Collins, Bruce
Veröffentlicht: (2025)
von: Collins, Bruce
Veröffentlicht: (2025)
Combining audio control and style transfer using latent diffusion
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
von: Demerlé, Nils, et al.
Veröffentlicht: (2024)
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
von: Lin, Tsung-En, et al.
Veröffentlicht: (2025)
von: Lin, Tsung-En, et al.
Veröffentlicht: (2025)
Adversarial multi-task underwater acoustic target recognition: towards robustness against various influential factors
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
A Dataset for Automatic Assessment of TTS Quality in Spanish
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
von: Welford, Alejandro Sosa, et al.
Veröffentlicht: (2025)
Mitigating data replication in text-to-audio generative diffusion models through anti-memorization guidance
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
von: Messina, Francisco, et al.
Veröffentlicht: (2025)
Sentiment analysis in non-fixed length audios using a Fully Convolutional Neural Network
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
von: García-Ordás, María Teresa, et al.
Veröffentlicht: (2024)
Decodable but not structured: linear probing enables Underwater Acoustic Target Recognition with pretrained audio embeddings
von: Hummel, Hilde I., et al.
Veröffentlicht: (2026)
von: Hummel, Hilde I., et al.
Veröffentlicht: (2026)
BELT-2: Bootstrapping EEG-to-Language representation alignment for multi-task brain decoding
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2024)
von: Zhou, Jinzhao, et al.
Veröffentlicht: (2024)
Recomposer: Event-roll-guided generative audio editing
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
Mixer is more than just a model
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
von: Ji, Qingfeng, et al.
Veröffentlicht: (2024)
Visual representations in the human brain are aligned with large language models
von: Doerig, Adrien, et al.
Veröffentlicht: (2022)
von: Doerig, Adrien, et al.
Veröffentlicht: (2022)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
von: Siuzdak, Hubert
Veröffentlicht: (2023)
von: Siuzdak, Hubert
Veröffentlicht: (2023)
Benchmarks and leaderboards for sound demixing tasks
von: Solovyev, Roman, et al.
Veröffentlicht: (2023)
von: Solovyev, Roman, et al.
Veröffentlicht: (2023)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
DistriBlock: Identifying adversarial audio samples by leveraging characteristics of the output distribution
von: Pizarro, Matías, et al.
Veröffentlicht: (2023)
von: Pizarro, Matías, et al.
Veröffentlicht: (2023)
Towards auditory attention decoding with noise-tagging: A pilot study
von: Scheppink, H. A., et al.
Veröffentlicht: (2024)
von: Scheppink, H. A., et al.
Veröffentlicht: (2024)
Joint sentiment analysis of lyrics and audio in music
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
Character-aware audio-visual subtitling in context
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
von: Huh, Jaesung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023) -
Fusion approaches for emotion recognition from speech using acoustic and text-based features
von: Pepino, Leonardo, et al.
Veröffentlicht: (2024) -
Training chord recognition models on artificially generated audio
von: Majchrzak, Martyna, et al.
Veröffentlicht: (2025) -
Transformation of audio embeddings into interpretable, concept-based representations
von: Zhang, Alice, et al.
Veröffentlicht: (2025) -
Benchmarking Time-localized Explanations for Audio Classification Models
von: Bolaños, Cecilia, et al.
Veröffentlicht: (2025)