Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Yifan, Han, Bing, Wang, Hui, Wang, Wei, Ma, Ziyang, Zhou, Long, Jin, Zengrui, Yang, Guanrou, Wang, Tianrui, Tan, Xu, Chen, Xie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Progressive Residual Extraction based Pre-training for Speech Representation Learning
por: Wang, Tianrui, et al.
Publicado: (2024)
por: Wang, Tianrui, et al.
Publicado: (2024)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
por: Wang, Huimeng, et al.
Publicado: (2024)
por: Wang, Huimeng, et al.
Publicado: (2024)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
por: Tu, Wenming, et al.
Publicado: (2025)
por: Tu, Wenming, et al.
Publicado: (2025)
Position: Towards Responsible Evaluation for Text-to-Speech
por: Yang, Yifan, et al.
Publicado: (2025)
por: Yang, Yifan, et al.
Publicado: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
k2SSL: A Faster and Better Framework for Self-Supervised Speech Representation Learning
por: Yang, Yifan, et al.
Publicado: (2024)
por: Yang, Yifan, et al.
Publicado: (2024)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
por: Song, Yakun, et al.
Publicado: (2024)
por: Song, Yakun, et al.
Publicado: (2024)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
por: Yang, Guanrou, et al.
Publicado: (2026)
por: Yang, Guanrou, et al.
Publicado: (2026)
Measuring Prosody Diversity in Zero-Shot TTS: A New Metric, Benchmark, and Exploration
por: Yang, Yifan, et al.
Publicado: (2025)
por: Yang, Yifan, et al.
Publicado: (2025)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
por: Yang, Guanrou, et al.
Publicado: (2025)
por: Yang, Guanrou, et al.
Publicado: (2025)
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
por: Yang, Xiaoyu, et al.
Publicado: (2025)
por: Yang, Xiaoyu, et al.
Publicado: (2025)
Augmenting Open-Vocabulary Dysarthric Speech Assessment with Human Perceptual Supervision
por: Jia, Kaimeng, et al.
Publicado: (2025)
por: Jia, Kaimeng, et al.
Publicado: (2025)
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
por: Wang, Tianrui, et al.
Publicado: (2026)
por: Wang, Tianrui, et al.
Publicado: (2026)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
por: Ma, Linhan, et al.
Publicado: (2025)
por: Ma, Linhan, et al.
Publicado: (2025)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
CTC-Assisted LLM-Based Contextual ASR
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
por: Li, Zhaoqing, et al.
Publicado: (2025)
por: Li, Zhaoqing, et al.
Publicado: (2025)
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
por: Li, Guojian, et al.
Publicado: (2026)
por: Li, Guojian, et al.
Publicado: (2026)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
por: Yang, Yifan, et al.
Publicado: (2024)
por: Yang, Yifan, et al.
Publicado: (2024)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
por: Liu, Wenrui, et al.
Publicado: (2025)
por: Liu, Wenrui, et al.
Publicado: (2025)
Advancing Multi-grained Alignment for Contrastive Language-Audio Pre-training
por: Li, Yiming, et al.
Publicado: (2024)
por: Li, Yiming, et al.
Publicado: (2024)
The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge
por: Guo, Yiwei, et al.
Publicado: (2024)
por: Guo, Yiwei, et al.
Publicado: (2024)
Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement
por: Wang, Tingting, et al.
Publicado: (2024)
por: Wang, Tingting, et al.
Publicado: (2024)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
por: Liu, Alexander H., et al.
Publicado: (2025)
por: Liu, Alexander H., et al.
Publicado: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
por: Du, Zhihao, et al.
Publicado: (2025)
por: Du, Zhihao, et al.
Publicado: (2025)
ClariCodec: Optimising Neural Speech Codes for 200bps Communication using Reinforcement Learning
por: Wang, Junyi, et al.
Publicado: (2026)
por: Wang, Junyi, et al.
Publicado: (2026)
Fine-Grained Quantitative Emotion Editing for Speech Generation
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
por: Li, Zhaoqing, et al.
Publicado: (2025)
por: Li, Zhaoqing, et al.
Publicado: (2025)
Learning Time-Graph Frequency Representation for Monaural Speech Enhancement
por: Wang, Tingting, et al.
Publicado: (2025)
por: Wang, Tingting, et al.
Publicado: (2025)
Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
por: Xiao, Yunchong, et al.
Publicado: (2026)
por: Xiao, Yunchong, et al.
Publicado: (2026)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
por: Shen, Siyuan, et al.
Publicado: (2024)
por: Shen, Siyuan, et al.
Publicado: (2024)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
por: Wang, Guansu, et al.
Publicado: (2025)
por: Wang, Guansu, et al.
Publicado: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
por: Zhou, Wangzixi, et al.
Publicado: (2026)
por: Zhou, Wangzixi, et al.
Publicado: (2026)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
por: Yang, Guanrou, et al.
Publicado: (2024)
por: Yang, Guanrou, et al.
Publicado: (2024)
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
por: Sun, Xingwei, et al.
Publicado: (2025)
por: Sun, Xingwei, et al.
Publicado: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
por: Du, Chenpeng, et al.
Publicado: (2024)
por: Du, Chenpeng, et al.
Publicado: (2024)
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling
por: Wang, Quanxiu, et al.
Publicado: (2024)
por: Wang, Quanxiu, et al.
Publicado: (2024)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
por: Jin, Zengrui, et al.
Publicado: (2022)
por: Jin, Zengrui, et al.
Publicado: (2022)
Ejemplares similares
-
Progressive Residual Extraction based Pre-training for Speech Representation Learning
por: Wang, Tianrui, et al.
Publicado: (2024) -
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
por: Wang, Huimeng, et al.
Publicado: (2024) -
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
por: Tu, Wenming, et al.
Publicado: (2025) -
Position: Towards Responsible Evaluation for Text-to-Speech
por: Yang, Yifan, et al.
Publicado: (2025) -
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)