Acoustic and perceptual differences between standard and accented speech and their voice clones
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Tianle, Sun, Chengzhe, Rose, Phil, Lyu, Siwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Forensic deepfake audio detection using segmental speech features
por: Yang, Tianle, et al.
Publicado: (2025)
por: Yang, Tianle, et al.
Publicado: (2025)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
People are poorly equipped to detect AI-powered voice clones
por: Barrington, Sarah, et al.
Publicado: (2024)
por: Barrington, Sarah, et al.
Publicado: (2024)
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
por: Mandai, Yuto, et al.
Publicado: (2025)
por: Mandai, Yuto, et al.
Publicado: (2025)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
por: Dietrich, Juergen
Publicado: (2026)
por: Dietrich, Juergen
Publicado: (2026)
Qualitative Approaches to Voice UX
por: Seaborn, Katie, et al.
Publicado: (2024)
por: Seaborn, Katie, et al.
Publicado: (2024)
Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
por: Fujii, Takao, et al.
Publicado: (2025)
por: Fujii, Takao, et al.
Publicado: (2025)
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
por: Rabin, Eyal, et al.
Publicado: (2025)
por: Rabin, Eyal, et al.
Publicado: (2025)
A Custom-Built Ambient Scribe Reduces Cognitive Load and Documentation Burden for Telehealth Clinicians
por: Morse, Justin, et al.
Publicado: (2025)
por: Morse, Justin, et al.
Publicado: (2025)
Arabic Little STT: Arabic Children Speech Recognition Dataset
por: Alkadri, Mouhand, et al.
Publicado: (2025)
por: Alkadri, Mouhand, et al.
Publicado: (2025)
Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
por: Choudhury, Ritabrata Roy
Publicado: (2024)
por: Choudhury, Ritabrata Roy
Publicado: (2024)
More-than-Human Storytelling: Designing Longitudinal Narrative Engagements with Generative AI
por: Fabre, Émilie, et al.
Publicado: (2025)
por: Fabre, Émilie, et al.
Publicado: (2025)
Step-Audio-EditX Technical Report
por: Yan, Chao, et al.
Publicado: (2025)
por: Yan, Chao, et al.
Publicado: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
por: Huang, Ailin, et al.
Publicado: (2025)
por: Huang, Ailin, et al.
Publicado: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
por: Chen, Qian, et al.
Publicado: (2025)
por: Chen, Qian, et al.
Publicado: (2025)
TouchAI: Exploring human-AI perceptual alignment in touch through language model representations
por: Zhong, Shu, et al.
Publicado: (2024)
por: Zhong, Shu, et al.
Publicado: (2024)
REALM: A Dataset of Real-World LLM Use Cases
por: Cheng, Jingwen, et al.
Publicado: (2025)
por: Cheng, Jingwen, et al.
Publicado: (2025)
TeachMaster: Generative Teaching via Code
por: Wang, Yuheng, et al.
Publicado: (2025)
por: Wang, Yuheng, et al.
Publicado: (2025)
Exploring Communication Strategies for Collaborative LLM Agents in Mathematical Problem-Solving
por: Zhang, Liang, et al.
Publicado: (2025)
por: Zhang, Liang, et al.
Publicado: (2025)
Social Welfare Function Leaderboard: When LLM Agents Allocate Social Welfare
por: Shi, Zhengliang, et al.
Publicado: (2025)
por: Shi, Zhengliang, et al.
Publicado: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
Language Model Can Listen While Speaking
por: Ma, Ziyang, et al.
Publicado: (2024)
por: Ma, Ziyang, et al.
Publicado: (2024)
EmoKnob: Enhance Voice Cloning with Fine-Grained Emotion Control
por: Chen, Haozhe, et al.
Publicado: (2024)
por: Chen, Haozhe, et al.
Publicado: (2024)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
por: Jiang, Xilin, et al.
Publicado: (2025)
por: Jiang, Xilin, et al.
Publicado: (2025)
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
por: Wang, Ke, et al.
Publicado: (2025)
por: Wang, Ke, et al.
Publicado: (2025)
Echoes of Humanity: Exploring the Perceived Humanness of AI Music
por: Figueiredo, Flavio, et al.
Publicado: (2025)
por: Figueiredo, Flavio, et al.
Publicado: (2025)
An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
por: Chen, Sheng-Kai, et al.
Publicado: (2026)
por: Chen, Sheng-Kai, et al.
Publicado: (2026)
Same Words, Different Judgments: How Preferences Vary Across Modalities
por: Broukhim, Aaron, et al.
Publicado: (2026)
por: Broukhim, Aaron, et al.
Publicado: (2026)
BREATH: A Bio-Radar Embodied Agent for Tonal and Human-Aware Diffusion Music Generation
por: Wang, Yunzhe, et al.
Publicado: (2025)
por: Wang, Yunzhe, et al.
Publicado: (2025)
Opening Musical Creativity? Embedded Ideologies in Generative-AI Music Systems
por: Pram, Liam, et al.
Publicado: (2025)
por: Pram, Liam, et al.
Publicado: (2025)
The Ghost in the Keys: A Disklavier Demo for Human-AI Musical Co-Creativity
por: Bradshaw, Louis, et al.
Publicado: (2025)
por: Bradshaw, Louis, et al.
Publicado: (2025)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
por: Morreale, Fabio, et al.
Publicado: (2025)
por: Morreale, Fabio, et al.
Publicado: (2025)
Human-Centred LLM Privacy Audits: Findings and Frictions
por: Staufer, Dimitri, et al.
Publicado: (2026)
por: Staufer, Dimitri, et al.
Publicado: (2026)
From Binary to Bilingual: How the National Weather Service is Using Artificial Intelligence to Develop a Comprehensive Translation Program
por: Trujillo-Falcon, Joseph E., et al.
Publicado: (2025)
por: Trujillo-Falcon, Joseph E., et al.
Publicado: (2025)
An Empirical Investigation of Gender Stereotype Representation in Large Language Models: The Italian Case
por: Giachino, Gioele, et al.
Publicado: (2025)
por: Giachino, Gioele, et al.
Publicado: (2025)
A perishable ability? The future of writing in the face of generative artificial intelligence
por: Cunha, Evandro L. T. P.
Publicado: (2025)
por: Cunha, Evandro L. T. P.
Publicado: (2025)
A validity-guided workflow for robust large language model research in psychology
por: Lin, Zhicheng
Publicado: (2025)
por: Lin, Zhicheng
Publicado: (2025)
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
por: Lin, Zhicheng
Publicado: (2025)
por: Lin, Zhicheng
Publicado: (2025)
Conversational DNA: A New Visual Language for Understanding Dialogue Structure in Human and AI
por: Lin, Baihan
Publicado: (2025)
por: Lin, Baihan
Publicado: (2025)
Who's Asking? Simulating Role-Based Questions for Conversational AI Evaluation
por: Kaur, Navreet, et al.
Publicado: (2025)
por: Kaur, Navreet, et al.
Publicado: (2025)
Ejemplares similares
-
Forensic deepfake audio detection using segmental speech features
por: Yang, Tianle, et al.
Publicado: (2025) -
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
por: Yang, Tianle, et al.
Publicado: (2026) -
People are poorly equipped to detect AI-powered voice clones
por: Barrington, Sarah, et al.
Publicado: (2024) -
Super Kawaii Vocalics: Amplifying the "Cute" Factor in Computer Voice
por: Mandai, Yuto, et al.
Publicado: (2025) -
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
por: Dietrich, Juergen
Publicado: (2026)