A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Zhu, Gao, Xiyuan, Zhang, Yuqing, Nayak, Shekhar, Coler, Matt |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
por: Li, Zhu, et al.
Publicado: (2025)
por: Li, Zhu, et al.
Publicado: (2025)
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
por: Li, Zhu, et al.
Publicado: (2025)
por: Li, Zhu, et al.
Publicado: (2025)
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025)
PSST! Prosodic Speech Segmentation with Transformers
por: Roll, Nathan, et al.
Publicado: (2023)
por: Roll, Nathan, et al.
Publicado: (2023)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
por: Sanders, Nicholas, et al.
Publicado: (2025)
por: Sanders, Nicholas, et al.
Publicado: (2025)
Which Prosodic Features Matter Most for Pragmatics?
por: Ward, Nigel G., et al.
Publicado: (2024)
por: Ward, Nigel G., et al.
Publicado: (2024)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
por: Ohnaka, Hien, et al.
Publicado: (2025)
por: Ohnaka, Hien, et al.
Publicado: (2025)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
por: de Seyssel, Maureen, et al.
Publicado: (2023)
por: de Seyssel, Maureen, et al.
Publicado: (2023)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
por: Sun, Haitong, et al.
Publicado: (2026)
por: Sun, Haitong, et al.
Publicado: (2026)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
por: Amooie, Reihaneh, et al.
Publicado: (2025)
por: Amooie, Reihaneh, et al.
Publicado: (2025)
Visual Cues Support Robust Turn-taking Prediction in Noise
por: Russell, Sam O'Connor, et al.
Publicado: (2025)
por: Russell, Sam O'Connor, et al.
Publicado: (2025)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoning
por: Li, Zhu, et al.
Publicado: (2026)
por: Li, Zhu, et al.
Publicado: (2026)
Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features
por: Khanjani, Zahra, et al.
Publicado: (2024)
por: Khanjani, Zahra, et al.
Publicado: (2024)
Leveraging the Interplay Between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
por: Jeon, Yejin, et al.
Publicado: (2024)
por: Jeon, Yejin, et al.
Publicado: (2024)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
por: Lux, Florian, et al.
Publicado: (2024)
por: Lux, Florian, et al.
Publicado: (2024)
BoSS: Beyond-Semantic Speech
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
por: Du, Yichao, et al.
Publicado: (2024)
por: Du, Yichao, et al.
Publicado: (2024)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
por: Cheng, Luyao, et al.
Publicado: (2023)
por: Cheng, Luyao, et al.
Publicado: (2023)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
por: Chen, Weidong, et al.
Publicado: (2025)
por: Chen, Weidong, et al.
Publicado: (2025)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
SpeechTaxi: On Multilingual Semantic Speech Classification
por: Keller, Lennart, et al.
Publicado: (2024)
por: Keller, Lennart, et al.
Publicado: (2024)
Semantic enrichment towards efficient speech representations
por: Laperrière, Gaëlle, et al.
Publicado: (2023)
por: Laperrière, Gaëlle, et al.
Publicado: (2023)
Llama-VITS: Enhancing TTS Synthesis with Semantic Awareness
por: Feng, Xincan, et al.
Publicado: (2024)
por: Feng, Xincan, et al.
Publicado: (2024)
Enhancing the Robustness of Contextual ASR to Varying Biasing Information Volumes Through Purified Semantic Correlation Joint Modeling
por: Gu, Yue, et al.
Publicado: (2025)
por: Gu, Yue, et al.
Publicado: (2025)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
por: Wang, Hankun, et al.
Publicado: (2024)
por: Wang, Hankun, et al.
Publicado: (2024)
Advancing Topic Segmentation of Broadcasted Speech with Multilingual Semantic Embeddings
por: Shukla, Sakshi Deo, et al.
Publicado: (2024)
por: Shukla, Sakshi Deo, et al.
Publicado: (2024)
AutoProsody: A Prosodic Feature Extraction Tool for Indian Languages
por: Thinakaran, Preethi, et al.
Publicado: (2025)
por: Thinakaran, Preethi, et al.
Publicado: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
por: Chary, Podakanti Satyajith
Publicado: (2024)
por: Chary, Podakanti Satyajith
Publicado: (2024)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
por: Manakul, Potsawee, et al.
Publicado: (2026)
por: Manakul, Potsawee, et al.
Publicado: (2026)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
por: Wang, Dingdong, et al.
Publicado: (2024)
por: Wang, Dingdong, et al.
Publicado: (2024)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
por: Gao, Yingying, et al.
Publicado: (2024)
por: Gao, Yingying, et al.
Publicado: (2024)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
por: Sun, Haiyang, et al.
Publicado: (2023)
por: Sun, Haiyang, et al.
Publicado: (2023)
Not that Groove: Zero-Shot Symbolic Music Editing
por: Zhang, Li
Publicado: (2025)
por: Zhang, Li
Publicado: (2025)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
por: Zhu, Han, et al.
Publicado: (2024)
por: Zhu, Han, et al.
Publicado: (2024)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
por: Xue, Jinlong, et al.
Publicado: (2024)
por: Xue, Jinlong, et al.
Publicado: (2024)
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
por: Wang, Chengyou, et al.
Publicado: (2026)
por: Wang, Chengyou, et al.
Publicado: (2026)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
por: Wang, Jianjin, et al.
Publicado: (2025)
por: Wang, Jianjin, et al.
Publicado: (2025)
Configurable Multilingual ASR with Speech Summary Representations
por: Zhu, Harrison, et al.
Publicado: (2024)
por: Zhu, Harrison, et al.
Publicado: (2024)
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
por: Yang, Runyan, et al.
Publicado: (2025)
por: Yang, Runyan, et al.
Publicado: (2025)
Ejemplares similares
-
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
por: Li, Zhu, et al.
Publicado: (2025) -
Leveraging Large Language Models for Sarcastic Speech Annotation in Sarcasm Detection
por: Li, Zhu, et al.
Publicado: (2025) -
Overcoming Data Scarcity in Multi-Dialectal Arabic ASR via Whisper Fine-Tuning
por: Özyilmaz, Ömer Tarik, et al.
Publicado: (2025) -
PSST! Prosodic Speech Segmentation with Transformers
por: Roll, Nathan, et al.
Publicado: (2023) -
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
por: Sanders, Nicholas, et al.
Publicado: (2025)