SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
Fuente:
arXiv
Salvato in:
| Autori principali: | Vallés-Pérez, Ivan, Beringer, Grzegorz, Bilinski, Piotr, Cook, Gary, Barra-Chicote, Roberto |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024)
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
di: Ma, Yi, et al.
Pubblicazione: (2025)
di: Ma, Yi, et al.
Pubblicazione: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
di: Cao, Yubing, et al.
Pubblicazione: (2025)
di: Cao, Yubing, et al.
Pubblicazione: (2025)
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
di: Hummel, Hilde I., et al.
Pubblicazione: (2025)
di: Hummel, Hilde I., et al.
Pubblicazione: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026)
di: Han, Seungu, et al.
Pubblicazione: (2026)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
di: Gong, Yitian, et al.
Pubblicazione: (2025)
di: Gong, Yitian, et al.
Pubblicazione: (2025)
Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
di: Ye, ShuQi, et al.
Pubblicazione: (2024)
di: Ye, ShuQi, et al.
Pubblicazione: (2024)
Investigating self-supervised features for expressive, multilingual voice conversion
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2025)
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2025)
Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
di: Dixit, Satvik, et al.
Pubblicazione: (2024)
Probing the Information Encoded in Neural-based Acoustic Models of Automatic Speech Recognition Systems
di: Raymondaud, Quentin, et al.
Pubblicazione: (2024)
di: Raymondaud, Quentin, et al.
Pubblicazione: (2024)
Contrastive Augmentation: An Unsupervised Learning Approach for Keyword Spotting in Speech Technology
di: Dai, Weinan, et al.
Pubblicazione: (2024)
di: Dai, Weinan, et al.
Pubblicazione: (2024)
A Small-footprint Acoustic Echo Cancellation Solution for Mobile Full-Duplex Speech Interactions
di: Jiang, Yiheng, et al.
Pubblicazione: (2025)
di: Jiang, Yiheng, et al.
Pubblicazione: (2025)
Cross-Corpus Validation of Speech Emotion Recognition in Urdu using Domain-Knowledge Acoustic Features
di: Talpur, Unzela, et al.
Pubblicazione: (2025)
di: Talpur, Unzela, et al.
Pubblicazione: (2025)
Color-based Emotion Representation for Speech Emotion Recognition
di: Nagase, Ryotaro, et al.
Pubblicazione: (2026)
di: Nagase, Ryotaro, et al.
Pubblicazione: (2026)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Real-Time Speech Enhancement via a Hybrid ViT: A Dual-Input Acoustic-Image Feature Fusion
di: Bahmei, Behnaz, et al.
Pubblicazione: (2025)
di: Bahmei, Behnaz, et al.
Pubblicazione: (2025)
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
di: Ahn, Hoseong, et al.
Pubblicazione: (2026)
Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
MATER: Multi-level Acoustic and Textual Emotion Representation for Interpretable Speech Emotion Recognition
di: Jon, Hyo Jin, et al.
Pubblicazione: (2025)
di: Jon, Hyo Jin, et al.
Pubblicazione: (2025)
Meta-Learning Approaches for Improving Detection of Unseen Speech Deepfakes
di: Kukanov, Ivan, et al.
Pubblicazione: (2024)
di: Kukanov, Ivan, et al.
Pubblicazione: (2024)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
di: Chu, Yunji, et al.
Pubblicazione: (2024)
di: Chu, Yunji, et al.
Pubblicazione: (2024)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
di: Jiang, Xue, et al.
Pubblicazione: (2025)
di: Jiang, Xue, et al.
Pubblicazione: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
di: Hwang, Injune, et al.
Pubblicazione: (2024)
di: Hwang, Injune, et al.
Pubblicazione: (2024)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
di: Har-Tuv, Nadav, et al.
Pubblicazione: (2025)
di: Har-Tuv, Nadav, et al.
Pubblicazione: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
di: Erol, Mehmet Hamza, et al.
Pubblicazione: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
di: Cao, Tianyu, et al.
Pubblicazione: (2026)
di: Cao, Tianyu, et al.
Pubblicazione: (2026)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
di: Zhang, Alice, et al.
Pubblicazione: (2025)
di: Zhang, Alice, et al.
Pubblicazione: (2025)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
di: Li, Bohan, et al.
Pubblicazione: (2024)
di: Li, Bohan, et al.
Pubblicazione: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
di: Sun, Yujia, et al.
Pubblicazione: (2024)
di: Sun, Yujia, et al.
Pubblicazione: (2024)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
di: Wu, Tung-Yu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024) -
Enhancing the Stability of LLM-based Speech Generation Systems through Self-Supervised Representations
di: Martín-Cortinas, Álvaro, et al.
Pubblicazione: (2024) -
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
di: Ma, Yi, et al.
Pubblicazione: (2025) -
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
di: Cao, Yubing, et al.
Pubblicazione: (2025) -
The Computation of Generalized Embeddings for Underwater Acoustic Target Recognition using Contrastive Learning
di: Hummel, Hilde I., et al.
Pubblicazione: (2025)