Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Wonjune, Jia, Junteng, Wu, Chunyang, Zhou, Wei, Lakomkin, Egor, Gaur, Yashesh, Sari, Leda, Kim, Suyoun, Li, Ke, Mahadeokar, Jay, Kalinli, Ozlem |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
di: Zhou, Wei, et al.
Pubblicazione: (2024)
di: Zhou, Wei, et al.
Pubblicazione: (2024)
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024)
di: Jia, Junteng, et al.
Pubblicazione: (2024)
Faster Speech-LLaMA Inference with Multi-token Prediction
di: Raj, Desh, et al.
Pubblicazione: (2024)
di: Raj, Desh, et al.
Pubblicazione: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
di: Xie, Jiamin, et al.
Pubblicazione: (2023)
di: Xie, Jiamin, et al.
Pubblicazione: (2023)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
Can Speech LLMs Think while Listening?
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2025)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
Towards Machine Unlearning for Paralinguistic Speech Processing
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2025)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
di: Ma, Yingyi, et al.
Pubblicazione: (2024)
di: Ma, Yingyi, et al.
Pubblicazione: (2024)
Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
di: Girish, et al.
Pubblicazione: (2025)
di: Girish, et al.
Pubblicazione: (2025)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
di: Kim, Heeseung, et al.
Pubblicazione: (2024)
di: Kim, Heeseung, et al.
Pubblicazione: (2024)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
di: Jiang, Feng, et al.
Pubblicazione: (2025)
di: Jiang, Feng, et al.
Pubblicazione: (2025)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
di: Keren, Gil, et al.
Pubblicazione: (2024)
di: Keren, Gil, et al.
Pubblicazione: (2024)
Conversational Speech Naturalness Predictor
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling
di: Cheng, Yifan, et al.
Pubblicazione: (2025)
di: Cheng, Yifan, et al.
Pubblicazione: (2025)
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
di: Lou, Haowei, et al.
Pubblicazione: (2026)
di: Lou, Haowei, et al.
Pubblicazione: (2026)
SoulX-Podcast: Towards Realistic Long-form Podcasts with Dialectal and Paralinguistic Diversity
di: Xie, Hanke, et al.
Pubblicazione: (2025)
di: Xie, Hanke, et al.
Pubblicazione: (2025)
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
Recovering Performance in Speech Emotion Recognition from Discrete Tokens via Multi-Layer Fusion and Paralinguistic Feature Integration
di: Sun, Esther, et al.
Pubblicazione: (2026)
di: Sun, Esther, et al.
Pubblicazione: (2026)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
di: Sanders, Nicholas, et al.
Pubblicazione: (2025)
di: Sanders, Nicholas, et al.
Pubblicazione: (2025)
Bayesian Speech Synthesizers Can Learn from Multiple Teachers
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
di: Zhang, Ziyang, et al.
Pubblicazione: (2025)
Can LLMs Help Localize Fake Words in Partially Fake Speech?
di: Zhang, Lin, et al.
Pubblicazione: (2026)
di: Zhang, Lin, et al.
Pubblicazione: (2026)
Towards measuring fairness in speech recognition: Fair-Speech dataset
di: Veliche, Irina-Elena, et al.
Pubblicazione: (2024)
di: Veliche, Irina-Elena, et al.
Pubblicazione: (2024)
Resurfacing Paralinguistic Awareness in Large Audio Language Models
di: Yang, Hao, et al.
Pubblicazione: (2026)
di: Yang, Hao, et al.
Pubblicazione: (2026)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
Adapting WavLM for Speech Emotion Recognition
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
di: Latif, Siddique, et al.
Pubblicazione: (2023)
di: Latif, Siddique, et al.
Pubblicazione: (2023)
MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
di: Yeh, Sung-Lin, et al.
Pubblicazione: (2026)
CDSD: Chinese Dysarthria Speech Database
di: Wang, Yan, et al.
Pubblicazione: (2023)
di: Wang, Yan, et al.
Pubblicazione: (2023)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
Leveraging LLM for Stuttering Speech: A Unified Architecture Bridging Recognition and Event Detection
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
Dual-View Predictive Diffusion: Lightweight Speech Enhancement via Spectrogram-Image Synergy
di: Xue, Ke, et al.
Pubblicazione: (2026)
di: Xue, Ke, et al.
Pubblicazione: (2026)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
di: Mattursun, Alimjan, et al.
Pubblicazione: (2025)
DroFiT: A Lightweight Band-fused Frequency Attention Toward Real-time UAV Speech Enhancement
di: Lee, Jeongmin, et al.
Pubblicazione: (2025)
di: Lee, Jeongmin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
di: Zhou, Wei, et al.
Pubblicazione: (2024) -
Efficient Streaming LLM for Speech Recognition
di: Jia, Junteng, et al.
Pubblicazione: (2024) -
Faster Speech-LLaMA Inference with Multi-token Prediction
di: Raj, Desh, et al.
Pubblicazione: (2024) -
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
di: Yang, Yufeng, et al.
Pubblicazione: (2024) -
Dynamic ASR Pathways: An Adaptive Masking Approach Towards Efficient Pruning of A Multilingual ASR Model
di: Xie, Jiamin, et al.
Pubblicazione: (2023)