SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vallés-Pérez, Ivan, Beringer, Grzegorz, Bilinski, Piotr, Cook, Gary, Barra-Chicote, Roberto |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2024)
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2024)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
von: Ma, Yi, et al.
Veröffentlicht: (2025)
von: Ma, Yi, et al.
Veröffentlicht: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
Investigating self-supervised features for expressive, multilingual voice conversion
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2025)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
von: Raymondaud, Quentin, et al.
Veröffentlicht: (2024)
Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
von: Dai, Weinan, et al.
Veröffentlicht: (2024)
von: Dai, Weinan, et al.
Veröffentlicht: (2024)
A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
von: Jiang, Yiheng, et al.
Veröffentlicht: (2025)
von: Jiang, Yiheng, et al.
Veröffentlicht: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
von: Talpur, Unzela, et al.
Veröffentlicht: (2025)
Color-based Emotion Representation for Speech Emotion Recognition
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
von: Nagase, Ryotaro, et al.
Veröffentlicht: (2026)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
von: Bahmei, Behnaz, et al.
Veröffentlicht: (2025)
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
von: Ahn, Hoseong, et al.
Veröffentlicht: (2026)
von: Ahn, Hoseong, et al.
Veröffentlicht: (2026)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Jon, Hyo Jin, et al.
Veröffentlicht: (2025)
Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
von: Kukanov, Ivan, et al.
Veröffentlicht: (2024)
von: Kukanov, Ivan, et al.
Veröffentlicht: (2024)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
von: Cao, Tianyu, et al.
Veröffentlicht: (2026)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2024)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024) -
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
von: Martín-Cortinas, Álvaro, et al.
Veröffentlicht: (2024) -
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
von: Ma, Yi, et al.
Veröffentlicht: (2025) -
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025) -
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
von: Hummel, Hilde I., et al.
Veröffentlicht: (2025)