Calliope: A TTS-based Narrated E-book Creator Ensuring Exact Synchronization, Privacy, and Layout Fidelity
Fuente:
arXiv
Saved in:
| Main Authors: | Hammer, Hugo L., Thambawita, Vajira, Halvorsen, Pål |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics
by: Riegler, Michael A., et al.
Published: (2025)
by: Riegler, Michael A., et al.
Published: (2025)
MOSS-TTS Technical Report
by: Gong, Yitian, et al.
Published: (2026)
by: Gong, Yitian, et al.
Published: (2026)
ECG-IMN: Interpretable Mesomorphic Neural Networks for 12-Lead Electrocardiogram Interpretation
by: Thambawita, Vajira, et al.
Published: (2026)
by: Thambawita, Vajira, et al.
Published: (2026)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
by: Yang, Tianle, et al.
Published: (2026)
by: Yang, Tianle, et al.
Published: (2026)
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)
Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models
by: Riera, Pablo, et al.
Published: (2026)
by: Riera, Pablo, et al.
Published: (2026)
Speak, Edit, Repeat: High-Fidelity Voice Editing and Zero-Shot TTS with Cross-Attentive Mamba
by: Mohammad, Baher, et al.
Published: (2025)
by: Mohammad, Baher, et al.
Published: (2025)
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
by: Salehi, Pegah, et al.
Published: (2024)
by: Salehi, Pegah, et al.
Published: (2024)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
by: Liu, Jiaxuan, et al.
Published: (2024)
by: Liu, Jiaxuan, et al.
Published: (2024)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
by: Lian, Jiachen, et al.
Published: (2022)
by: Lian, Jiachen, et al.
Published: (2022)
Generative Artificial Intelligence, Musical Heritage and the Construction of Peace Narratives: A Case Study in Mali
by: Coulibaly, Nouhoum, et al.
Published: (2026)
by: Coulibaly, Nouhoum, et al.
Published: (2026)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
by: Bataev, Vladimir, et al.
Published: (2025)
by: Bataev, Vladimir, et al.
Published: (2025)
A Comparative Study of Decoding Strategies in Medical Text Generation
by: Presacan, Oriana, et al.
Published: (2025)
by: Presacan, Oriana, et al.
Published: (2025)
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning
by: Sigurgeirsson, Atli, et al.
Published: (2025)
by: Sigurgeirsson, Atli, et al.
Published: (2025)
No Verifiable Reward for Prosody: Toward Preference-Guided Prosody Learning in TTS
by: Shin, Seungyoun, et al.
Published: (2025)
by: Shin, Seungyoun, et al.
Published: (2025)
TTS-1 Technical Report
by: Atamanenko, Oleg, et al.
Published: (2025)
by: Atamanenko, Oleg, et al.
Published: (2025)
A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data
by: Chou, Cheng-Kang, et al.
Published: (2025)
by: Chou, Cheng-Kang, et al.
Published: (2025)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
by: Ma, Ziyang, et al.
Published: (2023)
by: Ma, Ziyang, et al.
Published: (2023)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
by: Handoyo, Ahmad Alfani, et al.
Published: (2024)
by: Handoyo, Ahmad Alfani, et al.
Published: (2024)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
by: Zhou, Siyi, et al.
Published: (2025)
by: Zhou, Siyi, et al.
Published: (2025)
FMSD-TTS: Few-shot Multi-Speaker Multi-Dialect Text-to-Speech Synthesis for Ü-Tsang, Amdo and Kham Speech Dataset Generation
by: Liu, Yutong, et al.
Published: (2025)
by: Liu, Yutong, et al.
Published: (2025)
CIPHER: Conformer-based Inference of Phonemes from High-density EEG
by: Madishetty, Varshith
Published: (2026)
by: Madishetty, Varshith
Published: (2026)
Sing it, Narrate it: Quality Musical Lyrics Translation
by: Ye, Zhuorui, et al.
Published: (2024)
by: Ye, Zhuorui, et al.
Published: (2024)
A novel LSTM music generator based on the fractional time-frequency feature extraction
by: Ya, Li, et al.
Published: (2026)
by: Ya, Li, et al.
Published: (2026)
Data-efficient Targeted Token-level Preference Optimization for LLM-based Text-to-Speech
by: Kotoge, Rikuto, et al.
Published: (2025)
by: Kotoge, Rikuto, et al.
Published: (2025)
Medico 2025: Visual Question Answering for Gastrointestinal Imaging
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
by: Lee, Keon, et al.
Published: (2024)
by: Lee, Keon, et al.
Published: (2024)
InspireMusic: Integrating Super Resolution and Large Language Model for High-Fidelity Long-Form Music Generation
by: Zhang, Chong, et al.
Published: (2025)
by: Zhang, Chong, et al.
Published: (2025)
Regularized Federated Learning for Privacy-Preserving Dysarthric and Elderly Speech Recognition
by: Zhong, Tao, et al.
Published: (2025)
by: Zhong, Tao, et al.
Published: (2025)
A Critical Review of the Need for Knowledge-Centric Evaluation of Quranic Recitation
by: Al-Kharusi, Mohammed Hilal, et al.
Published: (2025)
by: Al-Kharusi, Mohammed Hilal, et al.
Published: (2025)
AI-Generated Song Detection via Lyrics Transcripts
by: Frohmann, Markus, et al.
Published: (2025)
by: Frohmann, Markus, et al.
Published: (2025)
Spatial Audio Motion Understanding and Reasoning
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
by: Sridhar, Arvind Krishna, et al.
Published: (2025)
Abusive music and song transformation using GenAI and LLMs
by: Choi, Jiyang, et al.
Published: (2026)
by: Choi, Jiyang, et al.
Published: (2026)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
by: Yang, Chenchen, et al.
Published: (2026)
by: Yang, Chenchen, et al.
Published: (2026)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
by: Lin, Tzu-Quan, et al.
Published: (2025)
by: Lin, Tzu-Quan, et al.
Published: (2025)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
by: Wang, Peng, et al.
Published: (2026)
by: Wang, Peng, et al.
Published: (2026)
Do LLM Decoders Listen Fairly? Benchmarking How Language Model Priors Shape Bias in Speech Recognition
by: Ginjala, Srishti, et al.
Published: (2026)
by: Ginjala, Srishti, et al.
Published: (2026)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition
by: Lee, Junseok, et al.
Published: (2026)
by: Lee, Junseok, et al.
Published: (2026)
Similar Items
-
Using Large Language Models to Suggest Informative Prior Distributions in Bayesian Statistics
by: Riegler, Michael A., et al.
Published: (2025) -
MOSS-TTS Technical Report
by: Gong, Yitian, et al.
Published: (2026) -
ECG-IMN: Interpretable Mesomorphic Neural Networks for 12-Lead Electrocardiogram Interpretation
by: Thambawita, Vajira, et al.
Published: (2026) -
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
by: Yang, Tianle, et al.
Published: (2026) -
A2TTS: TTS for Low Resource Indian Languages
by: Bhadoriya, Ayush Singh, et al.
Published: (2025)