Acoustic and perceptual differences between standard and accented speech and their voice clones
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Tianle, Sun, Chengzhe, Rose, Phil, Lyu, Siwei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forensic deepfake audio detection using segmental speech features
by: Yang, Tianle, et al.
Published: (2025)
by: Yang, Tianle, et al.
Published: (2025)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
by: Yang, Tianle, et al.
Published: (2026)
by: Yang, Tianle, et al.
Published: (2026)
People are poorly equipped to detect AI-powered voice clones
by: Barrington, Sarah, et al.
Published: (2024)
by: Barrington, Sarah, et al.
Published: (2024)
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
by: Mandai, Yuto, et al.
Published: (2025)
by: Mandai, Yuto, et al.
Published: (2025)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
by: Dietrich, Juergen
Published: (2026)
by: Dietrich, Juergen
Published: (2026)
Qualitative Approaches to Voice UX
by: Seaborn, Katie, et al.
Published: (2024)
by: Seaborn, Katie, et al.
Published: (2024)
Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
by: Fujii, Takao, et al.
Published: (2025)
by: Fujii, Takao, et al.
Published: (2025)
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
by: Rabin, Eyal, et al.
Published: (2025)
by: Rabin, Eyal, et al.
Published: (2025)
A Custom-Built Ambient Scribe Reduces Cognitive Load and Documentation Burden for Telehealth Clinicians
by: Morse, Justin, et al.
Published: (2025)
by: Morse, Justin, et al.
Published: (2025)
Arabic Little STT: Arabic Children Speech Recognition Dataset
by: Alkadri, Mouhand, et al.
Published: (2025)
by: Alkadri, Mouhand, et al.
Published: (2025)
Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
by: Choudhury, Ritabrata Roy
Published: (2024)
by: Choudhury, Ritabrata Roy
Published: (2024)
More-than-Human Storytelling: Designing Longitudinal Narrative Engagements with Generative AI
by: Fabre, Émilie, et al.
Published: (2025)
by: Fabre, Émilie, et al.
Published: (2025)
Step-Audio-EditX Technical Report
by: Yan, Chao, et al.
Published: (2025)
by: Yan, Chao, et al.
Published: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
by: Huang, Ailin, et al.
Published: (2025)
by: Huang, Ailin, et al.
Published: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
by: Chen, Qian, et al.
Published: (2025)
by: Chen, Qian, et al.
Published: (2025)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
by: Zhong, Shu, et al.
Published: (2024)
by: Zhong, Shu, et al.
Published: (2024)
REALM: A Dataset of Real-World LLM Use Cases
by: Cheng, Jingwen, et al.
Published: (2025)
by: Cheng, Jingwen, et al.
Published: (2025)
TeachMaster: Generative Teaching via Code
by: Wang, Yuheng, et al.
Published: (2025)
by: Wang, Yuheng, et al.
Published: (2025)
Exploring Communication Strategies for Collaborative LLM Agents in Mathematical Problem-Solving
by: Zhang, Liang, et al.
Published: (2025)
by: Zhang, Liang, et al.
Published: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
by: Shi, Zhengliang, et al.
Published: (2025)
by: Shi, Zhengliang, et al.
Published: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
by: Han, Zhichen, et al.
Published: (2024)
by: Han, Zhichen, et al.
Published: (2024)
Language Model Can Listen While Speaking
by: Ma, Ziyang, et al.
Published: (2024)
by: Ma, Ziyang, et al.
Published: (2024)
EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
by: Chen, Haozhe, et al.
Published: (2024)
by: Chen, Haozhe, et al.
Published: (2024)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
by: Jiang, Xilin, et al.
Published: (2025)
by: Jiang, Xilin, et al.
Published: (2025)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
by: Wang, Ke, et al.
Published: (2025)
by: Wang, Ke, et al.
Published: (2025)
Echoes of Humanity: Exploring the Perceived Humanness of AI Music
by: Figueiredo, Flavio, et al.
Published: (2025)
by: Figueiredo, Flavio, et al.
Published: (2025)
An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
by: Chen, Sheng-Kai, et al.
Published: (2026)
by: Chen, Sheng-Kai, et al.
Published: (2026)
Same Words, Different Judgments: How Preferences Vary Across Modalities
by: Broukhim, Aaron, et al.
Published: (2026)
by: Broukhim, Aaron, et al.
Published: (2026)
BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation
by: Wang, Yunzhe, et al.
Published: (2025)
by: Wang, Yunzhe, et al.
Published: (2025)
Opening Musical Creativity? Embedded Ideologies in Generative-AI Music Systems
by: Pram, Liam, et al.
Published: (2025)
by: Pram, Liam, et al.
Published: (2025)
The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
by: Bradshaw, Louis, et al.
Published: (2025)
by: Bradshaw, Louis, et al.
Published: (2025)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
by: Morreale, Fabio, et al.
Published: (2025)
by: Morreale, Fabio, et al.
Published: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
From Binary to Bilingual: How the National Weather Service is Using Artificial Intelligence to Develop a Comprehensive Translation Program
by: Trujillo-Falcon, Joseph E., et al.
Published: (2025)
by: Trujillo-Falcon, Joseph E., et al.
Published: (2025)
An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case
by: Giachino, Gioele, et al.
Published: (2025)
by: Giachino, Gioele, et al.
Published: (2025)
A perishable ability? The future of writing in the face of generative artificial intelligence
by: Cunha, Evandro L. T. P.
Published: (2025)
by: Cunha, Evandro L. T. P.
Published: (2025)
A validity-guided workflow for robust large language model research in psychology
by: Lin, Zhicheng
Published: (2025)
by: Lin, Zhicheng
Published: (2025)
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
by: Lin, Zhicheng
Published: (2025)
by: Lin, Zhicheng
Published: (2025)
Conversational DNA: A New Visual Language for Understanding Dialogue Structure in Human and AI
by: Lin, Baihan
Published: (2025)
by: Lin, Baihan
Published: (2025)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
by: Kaur, Navreet, et al.
Published: (2025)
by: Kaur, Navreet, et al.
Published: (2025)
Similar Items
-
Forensic deepfake audio detection using segmental speech features
by: Yang, Tianle, et al.
Published: (2025) -
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
by: Yang, Tianle, et al.
Published: (2026) -
People are poorly equipped to detect AI-powered voice clones
by: Barrington, Sarah, et al.
Published: (2024) -
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
by: Mandai, Yuto, et al.
Published: (2025) -
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
by: Dietrich, Juergen
Published: (2026)