Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhu, Zhang, Yuqing, Gao, Xiyuan, Nayak, Shekhar, Coler, Matt |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025)
von: Li, Zhu, et al.
Veröffentlicht: (2025)
A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
von: Li, Zhu, et al.
Veröffentlicht: (2024)
von: Li, Zhu, et al.
Veröffentlicht: (2024)
PSST! Prosodic Speech Segmentation with Transformers
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
von: Roll, Nathan, et al.
Veröffentlicht: (2023)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025)
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
von: Sun, Haitong, et al.
Veröffentlicht: (2026)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
von: Lux, Florian, et al.
Veröffentlicht: (2024)
von: Lux, Florian, et al.
Veröffentlicht: (2024)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
von: Özyilmaz, Ömer Tarik, et al.
Veröffentlicht: (2025)
SpeechTaxi: On Multilingual Semantic Speech Classification
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
von: Keller, Lennart, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
BoSS: Beyond-Semantic Speech
von: Wang, Qing, et al.
Veröffentlicht: (2025)
von: Wang, Qing, et al.
Veröffentlicht: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
von: Yang, Runyan, et al.
Veröffentlicht: (2024)
Next Tokens Denoising for Speech Synthesis
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
von: Liu, Yanqing, et al.
Veröffentlicht: (2025)
VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
von: Zheng, Zhisheng, et al.
Veröffentlicht: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
Generative Expressive Conversational Speech Synthesis
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Borderless Long Speech Synthesis
von: Song, Xingchen, et al.
Veröffentlicht: (2026)
von: Song, Xingchen, et al.
Veröffentlicht: (2026)
Boosting Large Language Model for Speech Synthesis: An Empirical Study
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
von: Hao, Hongkun, et al.
Veröffentlicht: (2023)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
von: Baas, Matthew, et al.
Veröffentlicht: (2023)
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
von: Du, Yichao, et al.
Veröffentlicht: (2024)
von: Du, Yichao, et al.
Veröffentlicht: (2024)
Autoregressive Speech Synthesis without Vector Quantization
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Which Prosodic Features Matter Most for Pragmatics?
von: Ward, Nigel G., et al.
Veröffentlicht: (2024)
von: Ward, Nigel G., et al.
Veröffentlicht: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
von: Si, Yuke, et al.
Veröffentlicht: (2025)
von: Si, Yuke, et al.
Veröffentlicht: (2025)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
von: Jiang, Feng, et al.
Veröffentlicht: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
von: Li, Zhu, et al.
Veröffentlicht: (2025) -
A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
von: Li, Zhu, et al.
Veröffentlicht: (2024) -
PSST! Prosodic Speech Segmentation with Transformers
von: Roll, Nathan, et al.
Veröffentlicht: (2023) -
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023) -
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
von: Ohnaka, Hien, et al.
Veröffentlicht: (2025)