Refining Self-Supervised Learnt Speech Representation using Brain Activations
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Hengyu, Mei, Kangdi, Liu, Zhaoci, Ai, Yang, Chen, Liping, Zhang, Jie, Ling, Zhenhua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
por: Zuo, Keying, et al.
Publicado: (2024)
por: Zuo, Keying, et al.
Publicado: (2024)
Adversarial speech for voice privacy protection from Personalized Speech generation
por: Chen, Shihao, et al.
Publicado: (2024)
por: Chen, Shihao, et al.
Publicado: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
por: Liu, Jiaxuan, et al.
Publicado: (2024)
por: Liu, Jiaxuan, et al.
Publicado: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024)
por: Ai, Yang, et al.
Publicado: (2024)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
por: Stahl, Benjamin, et al.
Publicado: (2025)
por: Stahl, Benjamin, et al.
Publicado: (2025)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Universal Preference-Score-based Pairwise Speech Quality Assessment
por: Shi, Yu-Fei, et al.
Publicado: (2025)
por: Shi, Yu-Fei, et al.
Publicado: (2025)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
por: Chiu, Aemon Yat Fei, et al.
Publicado: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
por: Sato, Hiroshi, et al.
Publicado: (2025)
por: Sato, Hiroshi, et al.
Publicado: (2025)
Explicit Estimation of Magnitude and Phase Spectra in Parallel for High-Quality Speech Enhancement
por: Lu, Ye-Xin, et al.
Publicado: (2023)
por: Lu, Ye-Xin, et al.
Publicado: (2023)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
por: Li, Jialu, et al.
Publicado: (2024)
por: Li, Jialu, et al.
Publicado: (2024)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
por: Li, Sirui, et al.
Publicado: (2025)
por: Li, Sirui, et al.
Publicado: (2025)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
por: Lu, Ye-Xin, et al.
Publicado: (2024)
por: Lu, Ye-Xin, et al.
Publicado: (2024)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
por: Farhadipour, Aref, et al.
Publicado: (2024)
por: Farhadipour, Aref, et al.
Publicado: (2024)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
por: Kutsakov, Aleksandr, et al.
Publicado: (2025)
por: Kutsakov, Aleksandr, et al.
Publicado: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
por: Du, Chenpeng, et al.
Publicado: (2022)
por: Du, Chenpeng, et al.
Publicado: (2022)
The CCF AATC 2025 Speech Restoration Challenge: A Retrospective
por: Zhang, Junan, et al.
Publicado: (2025)
por: Zhang, Junan, et al.
Publicado: (2025)
Vision-Integrated High-Quality Neural Speech Coding
por: Guo, Yao, et al.
Publicado: (2025)
por: Guo, Yao, et al.
Publicado: (2025)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
por: Wang, Linqin, et al.
Publicado: (2024)
por: Wang, Linqin, et al.
Publicado: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2024)
por: Chen, Yafeng, et al.
Publicado: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
por: Chen, Yafeng, et al.
Publicado: (2023)
por: Chen, Yafeng, et al.
Publicado: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
por: Wang, Shih-heng, et al.
Publicado: (2024)
por: Wang, Shih-heng, et al.
Publicado: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
por: Poncelet, Jakob, et al.
Publicado: (2021)
por: Poncelet, Jakob, et al.
Publicado: (2021)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
por: Maekaku, Takashi, et al.
Publicado: (2025)
por: Maekaku, Takashi, et al.
Publicado: (2025)
AC-Mix: Self-Supervised Adaptation for Low-Resource Automatic Speech Recognition using Agnostic Contrastive Mixup
por: Carvalho, Carlos, et al.
Publicado: (2024)
por: Carvalho, Carlos, et al.
Publicado: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
por: Yamashita, Natsuo, et al.
Publicado: (2024)
por: Yamashita, Natsuo, et al.
Publicado: (2024)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
por: Ogg, Mattson, et al.
Publicado: (2025)
por: Ogg, Mattson, et al.
Publicado: (2025)
Self-Supervised Learning of Spatial Acoustic Representation with Cross-Channel Signal Reconstruction and Multi-Channel Conformer
por: Yang, Bing, et al.
Publicado: (2023)
por: Yang, Bing, et al.
Publicado: (2023)
Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
por: Li, Ruiqi, et al.
Publicado: (2024)
por: Li, Ruiqi, et al.
Publicado: (2024)
Speaker Recognition Using Isomorphic Graph Attention Network Based Pooling on Self-Supervised Representation
por: Ge, Zirui, et al.
Publicado: (2023)
por: Ge, Zirui, et al.
Publicado: (2023)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
por: Park, Chanho, et al.
Publicado: (2023)
por: Park, Chanho, et al.
Publicado: (2023)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
por: Wilkins, Julia, et al.
Publicado: (2025)
por: Wilkins, Julia, et al.
Publicado: (2025)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
por: Wilkins, Julia, et al.
Publicado: (2024)
por: Wilkins, Julia, et al.
Publicado: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
por: Chung, Raymond
Publicado: (2026)
por: Chung, Raymond
Publicado: (2026)
Say More with Less: Variable-Frame-Rate Speech Tokenization via Adaptive Clustering and Implicit Duration Coding
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
por: Zheng, Rui-Chen, et al.
Publicado: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
por: Lu, Ye-Xin, et al.
Publicado: (2025)
por: Lu, Ye-Xin, et al.
Publicado: (2025)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
por: Vaessen, Nik, et al.
Publicado: (2024)
por: Vaessen, Nik, et al.
Publicado: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
por: Osakuade, Opeyemi, et al.
Publicado: (2024)
por: Osakuade, Opeyemi, et al.
Publicado: (2024)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Ejemplares similares
-
Geometry-Constrained EEG Channel Selection for Brain-Assisted Speech Enhancement
por: Zuo, Keying, et al.
Publicado: (2024) -
Adversarial speech for voice privacy protection from Personalized Speech generation
por: Chen, Shihao, et al.
Publicado: (2024) -
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
por: Liu, Jiaxuan, et al.
Publicado: (2024) -
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
por: Liu, Rui, et al.
Publicado: (2024) -
Low-Latency Neural Speech Phase Prediction based on Parallel Estimation Architecture and Anti-Wrapping Losses for Speech Generation Tasks
por: Ai, Yang, et al.
Publicado: (2024)