Not My Voice! A Taxonomy of Ethical and Safety Harms of Speech Generators
Fuente:
arXiv
Saved in:
| Main Authors: | Hutiri, Wiebke, Papakyriakopoulos, Oresiti, Xiang, Alice |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
by: Hutiri, Wiebke, et al.
Published: (2025)
by: Hutiri, Wiebke, et al.
Published: (2025)
As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
by: Hutiri, Wiebke, et al.
Published: (2024)
by: Hutiri, Wiebke, et al.
Published: (2024)
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025)
by: Patel, Tanvina, et al.
Published: (2025)
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
by: Lameris, Harm, et al.
Published: (2025)
by: Lameris, Harm, et al.
Published: (2025)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
by: Yang, Guanrou, et al.
Published: (2025)
by: Yang, Guanrou, et al.
Published: (2025)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
by: Le, Chenyang, et al.
Published: (2024)
by: Le, Chenyang, et al.
Published: (2024)
ArVoice: A Multi-Speaker Dataset for Arabic Speech Synthesis
by: Toyin, Hawau Olamide, et al.
Published: (2025)
by: Toyin, Hawau Olamide, et al.
Published: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
by: Fan, Xiaoran, et al.
Published: (2025)
by: Fan, Xiaoran, et al.
Published: (2025)
Pheme: Efficient and Conversational Speech Generation
by: Budzianowski, Paweł, et al.
Published: (2024)
by: Budzianowski, Paweł, et al.
Published: (2024)
Qualitative Approaches to Voice UX
by: Seaborn, Katie, et al.
Published: (2024)
by: Seaborn, Katie, et al.
Published: (2024)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
FairLENS: Assessing Fairness in Law Enforcement Speech Recognition
by: Wang, Yicheng, et al.
Published: (2024)
by: Wang, Yicheng, et al.
Published: (2024)
Phonology-Guided Speech-to-Speech Translation for African Languages
by: Ochieng, Peter, et al.
Published: (2024)
by: Ochieng, Peter, et al.
Published: (2024)
Streaming Speech-to-Text Translation with a SpeechLLM
by: Parcollet, Titouan, et al.
Published: (2026)
by: Parcollet, Titouan, et al.
Published: (2026)
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
by: Mandai, Yuto, et al.
Published: (2025)
by: Mandai, Yuto, et al.
Published: (2025)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation
by: Gkritzali, Evangelia, et al.
Published: (2024)
by: Gkritzali, Evangelia, et al.
Published: (2024)
VoiceBench: Benchmarking LLM-Based Voice Assistants
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
by: Wan, Zhen, et al.
Published: (2025)
by: Wan, Zhen, et al.
Published: (2025)
PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
Voice EHR: Introducing Multimodal Audio Data for Health
by: Anibal, James, et al.
Published: (2024)
by: Anibal, James, et al.
Published: (2024)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
by: Teleki, Maria, et al.
Published: (2025)
by: Teleki, Maria, et al.
Published: (2025)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
by: Ma, Zhengrui, et al.
Published: (2024)
by: Ma, Zhengrui, et al.
Published: (2024)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
by: Yang, Guanrou, et al.
Published: (2026)
by: Yang, Guanrou, et al.
Published: (2026)
Exploring In-Context Learning of Textless Speech Language Model for Speech Classification Tasks
by: Hsu, Ming-Hao, et al.
Published: (2023)
by: Hsu, Ming-Hao, et al.
Published: (2023)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
by: Huzaifah, Muhammad, et al.
Published: (2024)
by: Huzaifah, Muhammad, et al.
Published: (2024)
KAME: Tandem Architecture for Enhancing Knowledge in Real-Time Speech-to-Speech Conversational AI
by: Kuroki, So, et al.
Published: (2025)
by: Kuroki, So, et al.
Published: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
by: Djanibekov, Amirbek, et al.
Published: (2026)
by: Djanibekov, Amirbek, et al.
Published: (2026)
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
by: Wang, Qiongqiong, et al.
Published: (2025)
by: Wang, Qiongqiong, et al.
Published: (2025)
Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
by: Fujii, Takao, et al.
Published: (2025)
by: Fujii, Takao, et al.
Published: (2025)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
by: Li, Yinghao Aaron, et al.
Published: (2024)
by: Li, Yinghao Aaron, et al.
Published: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
by: Peng, Puyuan, et al.
Published: (2024)
by: Peng, Puyuan, et al.
Published: (2024)
WAXAL: A Large-Scale Multilingual African Language Speech Corpus
by: Diack, Abdoulaye, et al.
Published: (2026)
by: Diack, Abdoulaye, et al.
Published: (2026)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
SimClass: A Classroom Speech Dataset Generated via Game Engine Simulation For Automatic Speech Recognition Research
by: Attia, Ahmed Adel, et al.
Published: (2025)
by: Attia, Ahmed Adel, et al.
Published: (2025)
Voice Communication Analysis in Esports
by: Vinot, Aymeric, et al.
Published: (2024)
by: Vinot, Aymeric, et al.
Published: (2024)
Can Speech LLMs Think while Listening?
by: Shih, Yi-Jen, et al.
Published: (2025)
by: Shih, Yi-Jen, et al.
Published: (2025)
VibeVoice Technical Report
by: Peng, Zhiliang, et al.
Published: (2025)
by: Peng, Zhiliang, et al.
Published: (2025)
Brilla AI: AI Contestant for the National Science and Maths Quiz
by: Boateng, George, et al.
Published: (2024)
by: Boateng, George, et al.
Published: (2024)
Similar Items
-
TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation
by: Hutiri, Wiebke, et al.
Published: (2025) -
As Biased as You Measure: Methodological Pitfalls of Bias Evaluations in Speaker Verification Research
by: Hutiri, Wiebke, et al.
Published: (2024) -
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025) -
Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
by: Lameris, Harm, et al.
Published: (2025) -
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
by: Yang, Guanrou, et al.
Published: (2025)