Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Mengjie, Liu, Lianbo, Fujita, Yusuke, Shi, Hao, Gao, Yuan, Koshkin, Roman, Sudo, Yui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
von: Shi, Hao, et al.
Veröffentlicht: (2026)
von: Shi, Hao, et al.
Veröffentlicht: (2026)
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
von: Koshkin, Roman, et al.
Veröffentlicht: (2026)
von: Koshkin, Roman, et al.
Veröffentlicht: (2026)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
von: Fujita, Yusuke, et al.
Veröffentlicht: (2025)
von: Fujita, Yusuke, et al.
Veröffentlicht: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
DuplexCascade: Full-Duplex Speech-to-Speech Dialogue with VAD-Free Cascaded ASR-LLM-TTS Pipeline and Micro-Turn Optimization
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
von: Yang, Jianing, et al.
Veröffentlicht: (2026)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
von: Quang, Trung Nguyen, et al.
Veröffentlicht: (2026)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
von: Papi, Sara, et al.
Veröffentlicht: (2026)
von: Papi, Sara, et al.
Veröffentlicht: (2026)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
When Voice Matters: Evidence of Gender Disparity in Positional Bias of SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
von: Kim, Taesoo, et al.
Veröffentlicht: (2025)
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
AlignCap: Aligning Speech Emotion Captioning to Human Preferences
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
von: Liang, Ziqi, et al.
Veröffentlicht: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
von: Kando, Shunsuke, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Distilling LLM Semantic Priors into Encoder-Only Multi-Talker ASR with Talker-Count Routing
von: Shi, Hao, et al.
Veröffentlicht: (2026) -
Streaming Translation and Transcription Through Speech-to-Text Causal Alignment
von: Koshkin, Roman, et al.
Veröffentlicht: (2026) -
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025) -
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2025) -
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation
von: Fujita, Yusuke, et al.
Veröffentlicht: (2025)