Knowing What to Stress: A Discourse-Conditioned Text-to-Speech Benchmark
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Turetzky, Arnon, Dekel, Avihu, Aronowitz, Hagai, Hoory, Ron, Adi, Yossi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024)
von: Roth, Amit, et al.
Veröffentlicht: (2024)
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
Spoken question answering for visual queries
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
von: Shabtay, Nimrod, et al.
Veröffentlicht: (2025)
WHISTRESS: Enriching Transcriptions with Sentence Stress Detection
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Salmon: A Suite for Acoustic Language Model Evaluation
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
von: Maimon, Gallil, et al.
Veröffentlicht: (2024)
Slamming: Training a Speech Language Model on One GPU in a Day
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speech
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
von: Menta, Venkata Pushpak Teja
Veröffentlicht: (2026)
PROFASR-BENCH: A Benchmark for Context-Conditioned ASR in High-Stakes Professional Speech
von: Piskala, Deepak Babu
Veröffentlicht: (2025)
von: Piskala, Deepak Babu
Veröffentlicht: (2025)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
von: Xu, Shuhao, et al.
Veröffentlicht: (2026)
von: Xu, Shuhao, et al.
Veröffentlicht: (2026)
POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
von: Li, Xuanchen, et al.
Veröffentlicht: (2025)
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
PRiSM: Benchmarking Phone Realization in Speech Models
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2026)
PART: Progressive Alignment Representation Training for Multilingual Speech-To-Text with LLMs
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
von: Zhang, Pei, et al.
Veröffentlicht: (2025)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2025)
von: Minixhofer, Christoph, et al.
Veröffentlicht: (2025)
AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition
von: Awobade, Busayo, et al.
Veröffentlicht: (2026)
von: Awobade, Busayo, et al.
Veröffentlicht: (2026)
SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
von: Liu, Ruohan, et al.
Veröffentlicht: (2026)
RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
von: Matiyali, Neeraj, et al.
Veröffentlicht: (2025)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
von: Akarsh, Sai, et al.
Veröffentlicht: (2024)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation
von: Poon, Crystal Min Hui, et al.
Veröffentlicht: (2025)
von: Poon, Crystal Min Hui, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Cross-Attention is Half Explanation in Speech-to-Text Models
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Soundwave: Less is More for Speech-Text Alignment in LLMs
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025) -
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024) -
Speech Synthesis From Continuous Features Using Per-Token Latent Diffusion
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024) -
A Language Modeling Approach to Diacritic-Free Hebrew TTS
von: Roth, Amit, et al.
Veröffentlicht: (2024) -
Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS
von: Aronowitz, Hagai, et al.
Veröffentlicht: (2026)