Investigation for Relative Voice Impression Estimation
Fuente:
arXiv
Salvato in:
| Autori principali: | Fujita, Kenichi, Ijima, Yusuke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Voice Impression Control in Zero-Shot TTS
di: Fujita, Kenichi, et al.
Pubblicazione: (2025)
di: Fujita, Kenichi, et al.
Pubblicazione: (2025)
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Multi-interaction TTS toward professional recording reproduction
di: Kanagawa, Hiroki, et al.
Pubblicazione: (2025)
di: Kanagawa, Hiroki, et al.
Pubblicazione: (2025)
Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
di: Sekkat, Chloé, et al.
Pubblicazione: (2024)
di: Sekkat, Chloé, et al.
Pubblicazione: (2024)
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
di: Bogavelli, Tara, et al.
Pubblicazione: (2026)
di: Bogavelli, Tara, et al.
Pubblicazione: (2026)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
di: Sankar, Ashwin, et al.
Pubblicazione: (2024)
di: Sankar, Ashwin, et al.
Pubblicazione: (2024)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2023)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2023)
Evaluation of Google's Voice Recognition and Sentence Classification for Health Care Applications
di: Uddin, Majbah, et al.
Pubblicazione: (2024)
di: Uddin, Majbah, et al.
Pubblicazione: (2024)
Description-based Controllable Text-to-Speech with Cross-Lingual Voice Control
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
di: Yamamoto, Ryuichi, et al.
Pubblicazione: (2024)
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
di: Quamer, Waris, et al.
Pubblicazione: (2026)
di: Quamer, Waris, et al.
Pubblicazione: (2026)
Cross-lingual Text-To-Speech with Flow-based Voice Conversion for Improved Pronunciation
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2022)
di: Ellinas, Nikolaos, et al.
Pubblicazione: (2022)
Analyzing Multimodal Features of Spontaneous Voice Assistant Commands for Mild Cognitive Impairment Detection
di: Lin, Nana, et al.
Pubblicazione: (2024)
di: Lin, Nana, et al.
Pubblicazione: (2024)
Moonshine v2: Ergodic Streaming Encoder ASR for Latency-Critical Speech Applications
di: Kudlur, Manjunath, et al.
Pubblicazione: (2026)
di: Kudlur, Manjunath, et al.
Pubblicazione: (2026)
PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization
di: Banerjee, Adhiraj, et al.
Pubblicazione: (2026)
di: Banerjee, Adhiraj, et al.
Pubblicazione: (2026)
WavLink: Compact Audio-Text Embeddings with a Global Whisper Token
di: Kumar, Gokul Karthik, et al.
Pubblicazione: (2026)
di: Kumar, Gokul Karthik, et al.
Pubblicazione: (2026)
Huntington Disease Automatic Speech Recognition with Biomarker Supervision
di: Wang, Charles L., et al.
Pubblicazione: (2026)
di: Wang, Charles L., et al.
Pubblicazione: (2026)
RO-N3WS: Enhancing Generalization in Low-Resource ASR with Diverse Romanian Speech Benchmarks
di: Diaconu, Alexandra, et al.
Pubblicazione: (2026)
di: Diaconu, Alexandra, et al.
Pubblicazione: (2026)
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
di: Yang, Xusheng, et al.
Pubblicazione: (2025)
di: Yang, Xusheng, et al.
Pubblicazione: (2025)
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
di: Tian, Jinchuan, et al.
Pubblicazione: (2025)
Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
di: King, Evan, et al.
Pubblicazione: (2025)
di: King, Evan, et al.
Pubblicazione: (2025)
Assessing Factual Music Comprehension in Large Audio Language Models
di: Lin, Daniel Chenyu, et al.
Pubblicazione: (2025)
di: Lin, Daniel Chenyu, et al.
Pubblicazione: (2025)
Audio MultiChallenge: A Multi-Turn Evaluation of Spoken Dialogue Systems on Natural Human Interaction
di: Gosai, Advait, et al.
Pubblicazione: (2025)
di: Gosai, Advait, et al.
Pubblicazione: (2025)
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
di: Lin, Zhiyu, et al.
Pubblicazione: (2025)
di: Lin, Zhiyu, et al.
Pubblicazione: (2025)
Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
di: Ferreira, Alef Iury Siqueira, et al.
Pubblicazione: (2025)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
di: Baali, Massa, et al.
Pubblicazione: (2025)
di: Baali, Massa, et al.
Pubblicazione: (2025)
Semantic Codebooks as Effective Priors for Neural Speech Compression
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
di: Bai, Liuyang, et al.
Pubblicazione: (2025)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
di: Peng, Puyuan, et al.
Pubblicazione: (2024)
Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
di: Platnick, Daniel, et al.
Pubblicazione: (2024)
di: Platnick, Daniel, et al.
Pubblicazione: (2024)
GTR-Voice: Articulatory Phonetics Informed Controllable Expressive Speech Synthesis
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
di: Li, Zehua Kcriss, et al.
Pubblicazione: (2024)
Label-Context-Dependent Internal Language Model Estimation for CTC
di: Yang, Zijian, et al.
Pubblicazione: (2025)
di: Yang, Zijian, et al.
Pubblicazione: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
di: Yang, Zijian, et al.
Pubblicazione: (2023)
di: Yang, Zijian, et al.
Pubblicazione: (2023)
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
di: Nan, Zheng, et al.
Pubblicazione: (2024)
di: Nan, Zheng, et al.
Pubblicazione: (2024)
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
di: Mohammad, Baher, et al.
Pubblicazione: (2025)
di: Mohammad, Baher, et al.
Pubblicazione: (2025)
Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
di: Munakata, Hokuto, et al.
Pubblicazione: (2024)
di: Munakata, Hokuto, et al.
Pubblicazione: (2024)
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
di: Sakai, Yusuke, et al.
Pubblicazione: (2024)
di: Sakai, Yusuke, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Voice Impression Control in Zero-Shot TTS
di: Fujita, Kenichi, et al.
Pubblicazione: (2025) -
Speech Rhythm-Based Speaker Embeddings Extraction from Phonemes and Phoneme Duration for Multi-Speaker Speech Synthesis
di: Fujita, Kenichi, et al.
Pubblicazione: (2024) -
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024) -
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024) -
Multi-interaction TTS toward professional recording reproduction
di: Kanagawa, Hiroki, et al.
Pubblicazione: (2025)