Spectrotemporal Modulation: Efficient and Interpretable Feature Representation for Classifying Speech, Music, and Environmental Sounds
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chang, Andrew, Li, Yike, Roman, Iran R., Poeppel, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
von: Roman, Iran R., et al.
Veröffentlicht: (2024)
von: Roman, Iran R., et al.
Veröffentlicht: (2024)
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024)
Focal Modulation Networks for Interpretable Sound Classification
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
von: Della Libera, Luca, et al.
Veröffentlicht: (2024)
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
von: Tjandra, Andros, et al.
Veröffentlicht: (2025)
von: Tjandra, Andros, et al.
Veröffentlicht: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Semantic-Aware Interpretable Multimodal Music Auto-Tagging
von: Patakis, Andreas, et al.
Veröffentlicht: (2025)
von: Patakis, Andreas, et al.
Veröffentlicht: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
von: Gaur, Neeraj, et al.
Veröffentlicht: (2024)
von: Gaur, Neeraj, et al.
Veröffentlicht: (2024)
Evaluating Disentangled Representations for Controllable Music Generation
von: Ibáñez-Martínez, Laura, et al.
Veröffentlicht: (2026)
von: Ibáñez-Martínez, Laura, et al.
Veröffentlicht: (2026)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
Discovering and Steering Interpretable Concepts in Large Generative Music Models
von: Singh, Nikhil, et al.
Veröffentlicht: (2025)
von: Singh, Nikhil, et al.
Veröffentlicht: (2025)
Learning Disentangled Speech Representations
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
RepCodec: A Speech Representation Codec for Speech Tokenization
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
Embedding-Space Diffusion for Zero-Shot Environmental Sound Classification
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
von: Sims, Ysobel, et al.
Veröffentlicht: (2024)
Challenges in Automated Processing of Speech from Child Wearables: The Case of Voice Type Classifier
von: Kunze, Tarek, et al.
Veröffentlicht: (2025)
von: Kunze, Tarek, et al.
Veröffentlicht: (2025)
Advanced Framework for Animal Sound Classification With Features Optimization
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
von: Yang, Qiang, et al.
Veröffentlicht: (2024)
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
von: Ciranni, Ruben, et al.
Veröffentlicht: (2024)
Interpreting Graphic Notation with MusicLDM: An AI Improvisation of Cornelius Cardew's Treatise
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
Feature Aggregation in Joint Sound Classification and Localization Neural Networks
von: Healy, Brendan, et al.
Veröffentlicht: (2023)
von: Healy, Brendan, et al.
Veröffentlicht: (2023)
Diffusion Synthesizer for Efficient Multilingual Speech to Speech Translation
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2024)
von: Hirschkind, Nameer, et al.
Veröffentlicht: (2024)
Guitar-TECHS: An Electric Guitar Dataset Covering Techniques, Musical Excerpts, Chords and Scales Using a Diverse Array of Hardware
von: Pedroza, Hegel, et al.
Veröffentlicht: (2025)
von: Pedroza, Hegel, et al.
Veröffentlicht: (2025)
Parameter-Efficient Transfer Learning for Music Foundation Models
von: Ding, Yiwei, et al.
Veröffentlicht: (2024)
von: Ding, Yiwei, et al.
Veröffentlicht: (2024)
Multiple Choice Learning for Efficient Speech Separation with Many Speakers
von: Perera, David, et al.
Veröffentlicht: (2024)
von: Perera, David, et al.
Veröffentlicht: (2024)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
von: Glazer, Neta, et al.
Veröffentlicht: (2025)
Anticipatory Music Transformer
von: Thickstun, John, et al.
Veröffentlicht: (2023)
von: Thickstun, John, et al.
Veröffentlicht: (2023)
Music Emotion Prediction Using Recurrent Neural Networks
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
von: Chang, Xinyu, et al.
Veröffentlicht: (2024)
Score-informed Music Source Separation: Improving Synthetic-to-real Generalization in Classical Music
von: Tunturi, Eetu, et al.
Veröffentlicht: (2025)
von: Tunturi, Eetu, et al.
Veröffentlicht: (2025)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
Feature Representations for Automatic Meerkat Vocalization Classification
von: Mahmoud, Imen Ben, et al.
Veröffentlicht: (2024)
von: Mahmoud, Imen Ben, et al.
Veröffentlicht: (2024)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
Online Symbolic Music Alignment with Offline Reinforcement Learning
von: Peter, Silvan David
Veröffentlicht: (2023)
von: Peter, Silvan David
Veröffentlicht: (2023)
Behind the Scenes: Mechanistic Interpretability of LoRA-adapted Whisper for Speech Emotion Recognition
von: Ma, Yujian, et al.
Veröffentlicht: (2025)
von: Ma, Yujian, et al.
Veröffentlicht: (2025)
Disentangling Textual and Acoustic Features of Neural Speech Representations
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
Cross-domain Sound Recognition for Efficient Underwater Data Analysis
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
von: Park, Jeongsoo, et al.
Veröffentlicht: (2023)
Adaptive Slimming for Scalable and Efficient Speech Enhancement
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
von: Miccini, Riccardo, et al.
Veröffentlicht: (2025)
Knowledge Distillation for Speech Denoising by Latent Representation Alignment with Cosine Distance
von: Luong, Diep, et al.
Veröffentlicht: (2025)
von: Luong, Diep, et al.
Veröffentlicht: (2025)
Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2025)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
von: Roman, Iran R., et al.
Veröffentlicht: (2024) -
Benchmarking Representations for Speech, Music, and Acoustic Events
von: La Quatra, Moreno, et al.
Veröffentlicht: (2024) -
Focal Modulation Networks for Interpretable Sound Classification
von: Della Libera, Luca, et al.
Veröffentlicht: (2024) -
Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
von: Tjandra, Andros, et al.
Veröffentlicht: (2025) -
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)