The Role of Prosodic and Lexical Cues in Turn-Taking with Self-Supervised Speech Representations
Fuente:
arXiv
Saved in:
| Main Authors: | Russell, Sam OConnor, Charuau, Delphine, Harte, Naomi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
by: Russell, Sam O'Connor, et al.
Published: (2025)
by: Russell, Sam O'Connor, et al.
Published: (2025)
Visual Cues Support Robust Turn-taking Prediction in Noise
by: Russell, Sam O'Connor, et al.
Published: (2025)
by: Russell, Sam O'Connor, et al.
Published: (2025)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
by: Wallbridge, Sarenne, et al.
Published: (2025)
by: Wallbridge, Sarenne, et al.
Published: (2025)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
by: Li, Zhu, et al.
Published: (2025)
by: Li, Zhu, et al.
Published: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
by: Sun, Haitong, et al.
Published: (2026)
by: Sun, Haitong, et al.
Published: (2026)
The Impact of Prosodic Segmentation on Speech Synthesis of Spontaneous Speech
by: Galdino, Julio Cesar, et al.
Published: (2025)
by: Galdino, Julio Cesar, et al.
Published: (2025)
Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
by: Shome, Debaditya, et al.
Published: (2023)
by: Shome, Debaditya, et al.
Published: (2023)
PSST! Prosodic Speech Segmentation with Transformers
by: Roll, Nathan, et al.
Published: (2023)
by: Roll, Nathan, et al.
Published: (2023)
A Functional Trade-off between Prosodic and Semantic Cues in Conveying Sarcasm
by: Li, Zhu, et al.
Published: (2024)
by: Li, Zhu, et al.
Published: (2024)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
by: Rajaa, Shangeth
Published: (2026)
by: Rajaa, Shangeth
Published: (2026)
RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis
by: Yang, Qing, et al.
Published: (2025)
by: Yang, Qing, et al.
Published: (2025)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
by: de Seyssel, Maureen, et al.
Published: (2023)
by: de Seyssel, Maureen, et al.
Published: (2023)
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
by: Wang, Yancheng, et al.
Published: (2026)
by: Wang, Yancheng, et al.
Published: (2026)
Modeling Turn-Taking with Semantically Informed Gestures
by: Suresh, Varsha, et al.
Published: (2025)
by: Suresh, Varsha, et al.
Published: (2025)
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications
by: Baes, Naomi, et al.
Published: (2024)
by: Baes, Naomi, et al.
Published: (2024)
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
by: He, Yinghui, et al.
Published: (2026)
by: He, Yinghui, et al.
Published: (2026)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
by: Ohnaka, Hien, et al.
Published: (2025)
by: Ohnaka, Hien, et al.
Published: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Interpreting the Role of Visemes in Audio-Visual Speech Recognition
by: Papadopoulos, Aristeidis, et al.
Published: (2025)
by: Papadopoulos, Aristeidis, et al.
Published: (2025)
Self-Supervised Speech Representations are More Phonetic than Semantic
by: Choi, Kwanghee, et al.
Published: (2024)
by: Choi, Kwanghee, et al.
Published: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
by: Osakuade, Opeyemi, et al.
Published: (2024)
by: Osakuade, Opeyemi, et al.
Published: (2024)
Prompt-Guided Turn-Taking Prediction
by: Inoue, Koji, et al.
Published: (2025)
by: Inoue, Koji, et al.
Published: (2025)
"Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking
by: Cavalcanti, Julio Cesar, et al.
Published: (2025)
by: Cavalcanti, Julio Cesar, et al.
Published: (2025)
Syn-TurnTurk: A Synthetic Dataset for Turn-Taking Prediction in Turkish Dialogues
by: Bayrak, Ahmet Tuğrul, et al.
Published: (2026)
by: Bayrak, Ahmet Tuğrul, et al.
Published: (2026)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
by: Qu, Jiaming, et al.
Published: (2025)
by: Qu, Jiaming, et al.
Published: (2025)
Scene Abstraction for Lexical Semantics: Structured Representations of Situated Meaning
by: Cho, Yejin, et al.
Published: (2026)
by: Cho, Yejin, et al.
Published: (2026)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
by: Li, Yuxin, et al.
Published: (2025)
by: Li, Yuxin, et al.
Published: (2025)
On the Contribution of Lexical Features to Speech Emotion Recognition
by: Combei, David
Published: (2025)
by: Combei, David
Published: (2025)
Enhancing Modern Supervised Word Sense Disambiguation Models by Semantic Lexical Resources
by: Melacci, Stefano, et al.
Published: (2024)
by: Melacci, Stefano, et al.
Published: (2024)
NaturalTurn: A Method to Segment Speech into Psychologically Meaningful Conversational Turns
by: Cooney, Gus, et al.
Published: (2024)
by: Cooney, Gus, et al.
Published: (2024)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
by: Hernandez, Abner, et al.
Published: (2026)
by: Hernandez, Abner, et al.
Published: (2026)
An Exploration of Mamba for Speech Self-Supervised Models
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
by: Park, Chanho, et al.
Published: (2023)
by: Park, Chanho, et al.
Published: (2023)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
by: Venkateswaran, Nitin, et al.
Published: (2025)
by: Venkateswaran, Nitin, et al.
Published: (2025)
Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
by: Hernandez, Abner, et al.
Published: (2026)
by: Hernandez, Abner, et al.
Published: (2026)
Where Do Self-Supervised Speech Models Become Unfair?
by: Herron, Felix, et al.
Published: (2026)
by: Herron, Felix, et al.
Published: (2026)
Similar Items
-
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
by: Russell, Sam O'Connor, et al.
Published: (2025) -
Visual Cues Support Robust Turn-taking Prediction in Noise
by: Russell, Sam O'Connor, et al.
Published: (2025) -
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025) -
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
by: Wallbridge, Sarenne, et al.
Published: (2025) -
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
by: Li, Zhu, et al.
Published: (2025)