EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Fuente:
arXiv
Guardado en:
| Autores principales: | de Seyssel, Maureen, D'Avirro, Antony, Williams, Adina, Dupoux, Emmanuel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
por: Poli, Maxime, et al.
Publicado: (2026)
por: Poli, Maxime, et al.
Publicado: (2026)
PSST! Prosodic Speech Segmentation with Transformers
por: Roll, Nathan, et al.
Publicado: (2023)
por: Roll, Nathan, et al.
Publicado: (2023)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
por: Li, Zhu, et al.
Publicado: (2025)
por: Li, Zhu, et al.
Publicado: (2025)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
por: Poli, Maxime, et al.
Publicado: (2024)
por: Poli, Maxime, et al.
Publicado: (2024)
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
por: Ohnaka, Hien, et al.
Publicado: (2025)
por: Ohnaka, Hien, et al.
Publicado: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
por: Chen, Weidong, et al.
Publicado: (2025)
por: Chen, Weidong, et al.
Publicado: (2025)
Prosodic ABX: A Language-Agnostic Method for Measuring Prosodic Contrast in Speech Representations
por: Sun, Haitong, et al.
Publicado: (2026)
por: Sun, Haitong, et al.
Publicado: (2026)
fastabx: A library for efficient computation of ABX discriminability
por: Poli, Maxime, et al.
Publicado: (2025)
por: Poli, Maxime, et al.
Publicado: (2025)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Which Evaluation for Which Model? A Taxonomy for Speech Model Assessment
por: de Seyssel, Maureen, et al.
Publicado: (2025)
por: de Seyssel, Maureen, et al.
Publicado: (2025)
CASPER: A Large Scale Spontaneous Speech Dataset
por: Xiao, Cihan, et al.
Publicado: (2025)
por: Xiao, Cihan, et al.
Publicado: (2025)
Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
por: Chary, Podakanti Satyajith
Publicado: (2024)
por: Chary, Podakanti Satyajith
Publicado: (2024)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
por: Jacobs, Christiaan, et al.
Publicado: (2025)
por: Jacobs, Christiaan, et al.
Publicado: (2025)
Hallucination Benchmark for Speech Foundation Models
por: Koudounas, Alkis, et al.
Publicado: (2025)
por: Koudounas, Alkis, et al.
Publicado: (2025)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
por: Akarsh, Sai, et al.
Publicado: (2024)
por: Akarsh, Sai, et al.
Publicado: (2024)
Textually Pretrained Speech Language Models
por: Hassid, Michael, et al.
Publicado: (2023)
por: Hassid, Michael, et al.
Publicado: (2023)
Unveiling Biases while Embracing Sustainability: Assessing the Dual Challenges of Automatic Speech Recognition Systems
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
por: Kulkarni, Ajinkya, et al.
Publicado: (2025)
Detecting the Undetectable: Assessing the Efficacy of Current Spoof Detection Methods Against Seamless Speech Edits
por: Huang, Sung-Feng, et al.
Publicado: (2025)
por: Huang, Sung-Feng, et al.
Publicado: (2025)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2025)
por: Peng, Junyi, et al.
Publicado: (2025)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
por: Fang, Yuanbo, et al.
Publicado: (2025)
por: Fang, Yuanbo, et al.
Publicado: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
por: Li, Haoxun, et al.
Publicado: (2025)
por: Li, Haoxun, et al.
Publicado: (2025)
Benchmarking Automatic Speech Recognition Models for African Languages
por: Nahabwe, Alvin, et al.
Publicado: (2025)
por: Nahabwe, Alvin, et al.
Publicado: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
por: Zhang, Xin, et al.
Publicado: (2023)
por: Zhang, Xin, et al.
Publicado: (2023)
STAB: Speech Tokenizer Assessment Benchmark
por: Vashishth, Shikhar, et al.
Publicado: (2024)
por: Vashishth, Shikhar, et al.
Publicado: (2024)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Which Prosodic Features Matter Most for Pragmatics?
por: Ward, Nigel G., et al.
Publicado: (2024)
por: Ward, Nigel G., et al.
Publicado: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
por: Fan, Ruchao, et al.
Publicado: (2024)
por: Fan, Ruchao, et al.
Publicado: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
por: Huang, Kuan-Po, et al.
Publicado: (2023)
por: Huang, Kuan-Po, et al.
Publicado: (2023)
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
por: Tian, Jinchuan, et al.
Publicado: (2025)
por: Tian, Jinchuan, et al.
Publicado: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2024)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2024)
What Do Speech Foundation Models Not Learn About Speech?
por: Waheed, Abdul, et al.
Publicado: (2024)
por: Waheed, Abdul, et al.
Publicado: (2024)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
por: Shi, Jiatong, et al.
Publicado: (2023)
por: Shi, Jiatong, et al.
Publicado: (2023)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
por: Sanders, Nicholas, et al.
Publicado: (2025)
por: Sanders, Nicholas, et al.
Publicado: (2025)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
por: Jiang, Feng, et al.
Publicado: (2025)
por: Jiang, Feng, et al.
Publicado: (2025)
Compact Speech Translation Models via Discrete Speech Units Pretraining
por: Lam, Tsz Kin, et al.
Publicado: (2024)
por: Lam, Tsz Kin, et al.
Publicado: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025)
por: Futami, Hayato, et al.
Publicado: (2025)
Analysis of Speech Temporal Dynamics in the Context of Speaker Verification and Voice Anonymization
por: Tomashenko, Natalia, et al.
Publicado: (2024)
por: Tomashenko, Natalia, et al.
Publicado: (2024)
Ejemplares similares
-
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
por: Poli, Maxime, et al.
Publicado: (2026) -
PSST! Prosodic Speech Segmentation with Transformers
por: Roll, Nathan, et al.
Publicado: (2023) -
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
por: Li, Zhu, et al.
Publicado: (2025) -
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
por: Poli, Maxime, et al.
Publicado: (2024) -
Grapheme-Coherent Phonemic and Prosodic Annotation of Speech by Implicit and Explicit Grapheme Conditioning
por: Ohnaka, Hien, et al.
Publicado: (2025)