SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
Fuente:
arXiv
Salvato in:
| Autori principali: | Lu, Jingyu, Wang, Yuhan, Zhuo, Fan, Cheng, Xize, Pan, Changhao, Pu, Xueyi, Chen, Yifu, Wen, Chenyuhao, Liang, Tianle, Zhao, Zhou |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
di: Ji, Shengpeng, et al.
Pubblicazione: (2025)
di: Ji, Shengpeng, et al.
Pubblicazione: (2025)
WavChat: A Survey of Spoken Dialogue Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
di: Liu, Jingwen, et al.
Pubblicazione: (2025)
di: Liu, Jingwen, et al.
Pubblicazione: (2025)
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
di: Li, Ruiqi, et al.
Pubblicazione: (2026)
di: Li, Ruiqi, et al.
Pubblicazione: (2026)
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
di: Lei, Ke, et al.
Pubblicazione: (2026)
di: Lei, Ke, et al.
Pubblicazione: (2026)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
di: Cheng, Xize, et al.
Pubblicazione: (2025)
di: Cheng, Xize, et al.
Pubblicazione: (2025)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
di: Chen, Yifu, et al.
Pubblicazione: (2025)
di: Chen, Yifu, et al.
Pubblicazione: (2025)
FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems
di: Peng, Yizhou, et al.
Pubblicazione: (2025)
di: Peng, Yizhou, et al.
Pubblicazione: (2025)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
di: Nakata, Wataru, et al.
Pubblicazione: (2024)
LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
di: Zhang, Hao, et al.
Pubblicazione: (2025)
di: Zhang, Hao, et al.
Pubblicazione: (2025)
LALM-as-a-Judge: Benchmarking Large Audio-Language Models for Safety Evaluation in Multi-Turn Spoken Dialogues
di: Ivry, Amir, et al.
Pubblicazione: (2026)
di: Ivry, Amir, et al.
Pubblicazione: (2026)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
di: Ao, Junyi, et al.
Pubblicazione: (2024)
di: Ao, Junyi, et al.
Pubblicazione: (2024)
Paralinguistics-Enhanced Large Language Modeling of Spoken Dialogue
di: Lin, Guan-Ting, et al.
Pubblicazione: (2023)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2023)
RLBR: Reinforcement Learning with Biasing Rewards for Contextual Speech Large Language Models
di: Ren, Bo, et al.
Pubblicazione: (2026)
di: Ren, Bo, et al.
Pubblicazione: (2026)
Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
di: Peng, Yizhou, et al.
Pubblicazione: (2026)
Literary and Colloquial Dialect Identification for Tamil using Acoustic Features
di: Nanmalar, M., et al.
Pubblicazione: (2024)
di: Nanmalar, M., et al.
Pubblicazione: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
di: Lin, Guan-Ting, et al.
Pubblicazione: (2025)
di: Lin, Guan-Ting, et al.
Pubblicazione: (2025)
Towards a Japanese Full-duplex Spoken Dialogue System
di: Ohashi, Atsumoto, et al.
Pubblicazione: (2025)
di: Ohashi, Atsumoto, et al.
Pubblicazione: (2025)
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
di: Udupa, Sathvik, et al.
Pubblicazione: (2025)
DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
di: Koudounas, Alkis, et al.
Pubblicazione: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
di: Lu, Haitian, et al.
Pubblicazione: (2025)
di: Lu, Haitian, et al.
Pubblicazione: (2025)
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge
di: Wang, Chengyou, et al.
Pubblicazione: (2026)
di: Wang, Chengyou, et al.
Pubblicazione: (2026)
From Scores to Preferences: Redefining MOS Benchmarking for Speech Quality Reward Modeling
di: Cao, Yifei, et al.
Pubblicazione: (2025)
di: Cao, Yifei, et al.
Pubblicazione: (2025)
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
Synthetic Singers: A Review of Deep-Learning-based Singing Voice Synthesis Approaches
di: Pan, Changhao, et al.
Pubblicazione: (2026)
di: Pan, Changhao, et al.
Pubblicazione: (2026)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
di: Wu, Junkai, et al.
Pubblicazione: (2024)
di: Wu, Junkai, et al.
Pubblicazione: (2024)
TiCo: Time-Controllable Spoken Dialogue Model
di: Chang, Kai-Wei, et al.
Pubblicazione: (2026)
di: Chang, Kai-Wei, et al.
Pubblicazione: (2026)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
di: Ren, Yong, et al.
Pubblicazione: (2026)
di: Ren, Yong, et al.
Pubblicazione: (2026)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
di: Tu, Wenming, et al.
Pubblicazione: (2025)
di: Tu, Wenming, et al.
Pubblicazione: (2025)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Acoustic and Semantic Modeling of Emotion in Spoken Language
di: Dutta, Soumya
Pubblicazione: (2026)
di: Dutta, Soumya
Pubblicazione: (2026)
Aligning Spoken Dialogue Models from User Interactions
di: Wu, Anne, et al.
Pubblicazione: (2025)
di: Wu, Anne, et al.
Pubblicazione: (2025)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
HiPPO: Exploring A Novel Hierarchical Pronunciation Assessment Approach for Spoken Languages
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
di: Ji, Shengpeng, et al.
Pubblicazione: (2025) -
WavChat: A Survey of Spoken Dialogue Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024) -
EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
di: Liu, Jingwen, et al.
Pubblicazione: (2025) -
SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
di: Li, Ruiqi, et al.
Pubblicazione: (2026) -
Semantic-Aware Interruption Detection in Spoken Dialogue Systems: Benchmark, Metric, and Model
di: Xia, Kangxiang, et al.
Pubblicazione: (2026)