GSRM: Generative Speech Reward Model for Speech RLHF
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Maohao, Jayashankar, Tejas, Hanna, Osama, Kanda, Naoyuki, Wang, Yancheng, Žmolíková, Kateřina, Xie, Ruiming, Moritz, Niko, Xu, Anfeng, Gaur, Yashesh, Wornell, Gregory, He, Qing, Wu, Jilong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Conversational Speech Naturalness Predictor
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024)
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024)
di: Seide, Frank, et al.
Pubblicazione: (2024)
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
di: Lin, Ju, et al.
Pubblicazione: (2024)
di: Lin, Ju, et al.
Pubblicazione: (2024)
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
di: Shen, Maohao, et al.
Pubblicazione: (2024)
di: Shen, Maohao, et al.
Pubblicazione: (2024)
Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
di: Kang, Wonjune, et al.
Pubblicazione: (2024)
di: Kang, Wonjune, et al.
Pubblicazione: (2024)
DiariST: Streaming Speech Translation with Speaker Diarization
di: Yang, Mu, et al.
Pubblicazione: (2023)
di: Yang, Mu, et al.
Pubblicazione: (2023)
Score-of-Mixture Training: Training One-Step Generative Models Made Simple via Score Estimation of Mixture Distributions
di: Jayashankar, Tejas, et al.
Pubblicazione: (2025)
di: Jayashankar, Tejas, et al.
Pubblicazione: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
di: Subramanian, Aswin Shanmugam, et al.
Pubblicazione: (2025)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
di: Moritz, Niko, et al.
Pubblicazione: (2024)
di: Moritz, Niko, et al.
Pubblicazione: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
di: Yang, Fei, et al.
Pubblicazione: (2026)
di: Yang, Fei, et al.
Pubblicazione: (2026)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
di: Sun, Haoqin, et al.
Pubblicazione: (2026)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
di: Li, Tao, et al.
Pubblicazione: (2025)
di: Li, Tao, et al.
Pubblicazione: (2025)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
di: Wang, Peidong, et al.
Pubblicazione: (2025)
di: Wang, Peidong, et al.
Pubblicazione: (2025)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
VowelPrompt: Hearing Speech Emotions from Text via Vowel-level Prosodic Augmentation
di: Wang, Yancheng, et al.
Pubblicazione: (2026)
di: Wang, Yancheng, et al.
Pubblicazione: (2026)
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2026)
di: Chen, Youjun, et al.
Pubblicazione: (2026)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
di: Liu, Zhanxun, et al.
Pubblicazione: (2025)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
di: Chen, Peikun, et al.
Pubblicazione: (2024)
di: Chen, Peikun, et al.
Pubblicazione: (2024)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
di: Feng, Tiantian, et al.
Pubblicazione: (2023)
Coding Speech through Vocal Tract Kinematics
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
di: Gaur, Neeraj, et al.
Pubblicazione: (2024)
di: Gaur, Neeraj, et al.
Pubblicazione: (2024)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
di: Shi, Jiacheng, et al.
Pubblicazione: (2026)
di: Shi, Jiacheng, et al.
Pubblicazione: (2026)
DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
di: Ghosh, Suhita, et al.
Pubblicazione: (2026)
di: Ghosh, Suhita, et al.
Pubblicazione: (2026)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
SpeechT: Findings of the First Mentorship in Speech Translation
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
di: Moslem, Yasmin, et al.
Pubblicazione: (2025)
An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
di: Song, Haoyu, et al.
Pubblicazione: (2025)
di: Song, Haoyu, et al.
Pubblicazione: (2025)
Forensic Similarity for Speech Deepfakes
di: Negroni, Viola, et al.
Pubblicazione: (2025)
di: Negroni, Viola, et al.
Pubblicazione: (2025)
EMG-to-Speech with Fewer Channels
di: Hwang, Injune, et al.
Pubblicazione: (2026)
di: Hwang, Injune, et al.
Pubblicazione: (2026)
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
di: Li, Longhao, et al.
Pubblicazione: (2025)
di: Li, Longhao, et al.
Pubblicazione: (2025)
DisSR: Disentangling Speech Representation for Degradation-Prior Guided Cross-Domain Speech Restoration
di: Liang, Ziqi, et al.
Pubblicazione: (2026)
di: Liang, Ziqi, et al.
Pubblicazione: (2026)
Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
di: Sun, Anchen, et al.
Pubblicazione: (2025)
di: Sun, Anchen, et al.
Pubblicazione: (2025)
Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
di: Baskar, Murali Karthick, et al.
Pubblicazione: (2024)
di: Baskar, Murali Karthick, et al.
Pubblicazione: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-based Neural Speech Codec
di: Liu, Fei, et al.
Pubblicazione: (2026)
di: Liu, Fei, et al.
Pubblicazione: (2026)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Conversational Speech Naturalness Predictor
di: Xu, Anfeng, et al.
Pubblicazione: (2026) -
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
di: Zhao, Jinzheng, et al.
Pubblicazione: (2024) -
Speech ReaLLM -- Real-time Streaming Speech Recognition with Multimodal LLMs by Teaching the Flow of Time
di: Seide, Frank, et al.
Pubblicazione: (2024) -
AGADIR: Towards Array-Geometry Agnostic Directional Speech Recognition
di: Lin, Ju, et al.
Pubblicazione: (2024) -
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
di: Shen, Maohao, et al.
Pubblicazione: (2024)