"Dyadosyncrasy", Idiosyncrasy and Demographic Factors in Turn-Taking
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cavalcanti, Julio Cesar, Skantze, Gabriel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multilingual Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
Prompt-Guided Turn-Taking Prediction
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
von: Inoue, Koji, et al.
Veröffentlicht: (2025)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
von: Rajaa, Shangeth
Veröffentlicht: (2026)
von: Rajaa, Shangeth
Veröffentlicht: (2026)
From Turn-Taking to Synchronous Dialogue: A Survey of Full-Duplex Spoken Language Models
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
von: Chen, Yuxuan, et al.
Veröffentlicht: (2025)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
von: Inoue, Koji, et al.
Veröffentlicht: (2024)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2025)
VoXtream2: Full-stream TTS with dynamic speaking rate control
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
von: Torgashov, Nikita, et al.
Veröffentlicht: (2026)
Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
von: Jeon, Hyunbae, et al.
Veröffentlicht: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2025)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
Visual Cues Support Robust Turn-taking Prediction in Noise
von: Russell, Sam O'Connor, et al.
Veröffentlicht: (2025)
von: Russell, Sam O'Connor, et al.
Veröffentlicht: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2026)
von: Wei, Sheng-Lun, et al.
Veröffentlicht: (2026)
Speaker-Reasoner: Scaling Interaction Turns and Reasoning Patterns for Timestamped Speaker-Attributed ASR
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
von: Lin, Zhennan, et al.
Veröffentlicht: (2026)
The Voice of Equity: A Systematic Evaluation of Bias Mitigation Techniques for Speech-Based Cognitive Impairment Detection Across Architectures and Demographics
von: Haghbin, Yasaman, et al.
Veröffentlicht: (2026)
von: Haghbin, Yasaman, et al.
Veröffentlicht: (2026)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
Explainable Speech Emotion Recognition: Weighted Attribute Fairness to Model Demographic Contributions to Social Bias
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
von: Ogunnubi, Tomisin, et al.
Veröffentlicht: (2026)
Factor-Conditioned Speaking-Style Captioning
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
von: Ando, Atsushi, et al.
Veröffentlicht: (2024)
Factorized RVQ-GAN For Disentangled Speech Tokenization
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
von: Khurana, Sameer, et al.
Veröffentlicht: (2025)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
von: Wang, Peng, et al.
Veröffentlicht: (2023)
von: Wang, Peng, et al.
Veröffentlicht: (2023)
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Turn-taking annotation for quantitative and qualitative analyses of conversation
von: Kelterer, Anneliese, et al.
Veröffentlicht: (2025)
von: Kelterer, Anneliese, et al.
Veröffentlicht: (2025)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
von: Veluri, Bandhav, et al.
Veröffentlicht: (2024)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
von: Sekkat, Chloé, et al.
Veröffentlicht: (2024)
MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
von: Tam, Zhi Rui, et al.
Veröffentlicht: (2025)
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
von: Lu, Ke-Han, et al.
Veröffentlicht: (2025)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2025)
Exploring Generative Error Correction for Dysarthric Speech Recognition
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
von: La Quatra, Moreno, et al.
Veröffentlicht: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
von: Kamper, Herman, et al.
Veröffentlicht: (2025)
von: Kamper, Herman, et al.
Veröffentlicht: (2025)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
von: Louw, Retief, et al.
Veröffentlicht: (2025)
von: Louw, Retief, et al.
Veröffentlicht: (2025)
Optimizing ASR for Catalan-Spanish Code-Switching: A Comparative Analysis of Methodologies
von: Mena, Carlos, et al.
Veröffentlicht: (2025)
von: Mena, Carlos, et al.
Veröffentlicht: (2025)
Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
von: Allbert, Rumi, et al.
Veröffentlicht: (2025)
von: Allbert, Rumi, et al.
Veröffentlicht: (2025)
TokenVerse++: Towards Flexible Multitask Learning with Dynamic Task Activation
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
von: Kumar, Shashi, et al.
Veröffentlicht: (2025)
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
von: Qian, Kaizhi, et al.
Veröffentlicht: (2025)
von: Qian, Kaizhi, et al.
Veröffentlicht: (2025)
SpeechCT-CLIP: Distilling Text-Image Knowledge to Speech for Voice-Native Multimodal CT Analysis
von: Buess, Lukas, et al.
Veröffentlicht: (2025)
von: Buess, Lukas, et al.
Veröffentlicht: (2025)
Lightweight Prompt Biasing for Contextualized End-to-End ASR Systems
von: Ren, Bo, et al.
Veröffentlicht: (2025)
von: Ren, Bo, et al.
Veröffentlicht: (2025)
Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
von: Corrêa, Pedro, et al.
Veröffentlicht: (2025)
MVP: Multi-source Voice Pathology detection
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multilingual Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024) -
Prompt-Guided Turn-Taking Prediction
von: Inoue, Koji, et al.
Veröffentlicht: (2025) -
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
von: Arora, Siddhant, et al.
Veröffentlicht: (2025) -
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
von: Inoue, Koji, et al.
Veröffentlicht: (2024) -
DualTurn: Learning Turn-Taking from Dual-Channel Generative Speech Pretraining
von: Rajaa, Shangeth
Veröffentlicht: (2026)