High-Fidelity Neural Phonetic Posteriorgrams
Fuente:
arXiv
Guardado en:
| Autores principales: | Churchwell, Cameron, Morrison, Max, Pardo, Bryan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Fine-Grained and Interpretable Neural Speech Editing
por: Morrison, Max, et al.
Publicado: (2024)
por: Morrison, Max, et al.
Publicado: (2024)
Cross-domain Neural Pitch and Periodicity Estimation
por: Morrison, Max, et al.
Publicado: (2023)
por: Morrison, Max, et al.
Publicado: (2023)
Combolutional Neural Networks
por: Churchwell, Cameron, et al.
Publicado: (2025)
por: Churchwell, Cameron, et al.
Publicado: (2025)
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
por: Li, Zirui, et al.
Publicado: (2025)
por: Li, Zirui, et al.
Publicado: (2025)
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
por: Ahn, Sunghwan, et al.
Publicado: (2024)
por: Ahn, Sunghwan, et al.
Publicado: (2024)
Phonetic Richness for Improved Automatic Speaker Verification
por: Klein, Nicholas, et al.
Publicado: (2024)
por: Klein, Nicholas, et al.
Publicado: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
por: Ma, Yi, et al.
Publicado: (2026)
por: Ma, Yi, et al.
Publicado: (2026)
Code Drift: Towards Idempotent Neural Audio Codecs
por: O'Reilly, Patrick, et al.
Publicado: (2024)
por: O'Reilly, Patrick, et al.
Publicado: (2024)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
por: Chodroff, Eleanor, et al.
Publicado: (2024)
por: Chodroff, Eleanor, et al.
Publicado: (2024)
Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling
por: Zhou, Xuanru, et al.
Publicado: (2025)
por: Zhou, Xuanru, et al.
Publicado: (2025)
HiFi-Glot: High-Fidelity Neural Formant Synthesis with Differentiable Resonant Filters
por: Gu, Yicheng, et al.
Publicado: (2024)
por: Gu, Yicheng, et al.
Publicado: (2024)
MusicHiFi: Fast High-Fidelity Stereo Vocoding
por: Zhu, Ge, et al.
Publicado: (2024)
por: Zhu, Ge, et al.
Publicado: (2024)
Quantifying and Reducing Speaker Heterogeneity within the Common Voice Corpus for Phonetic Analysis
por: Zhang, Miao, et al.
Publicado: (2025)
por: Zhang, Miao, et al.
Publicado: (2025)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
por: Li, Jialu, et al.
Publicado: (2023)
por: Li, Jialu, et al.
Publicado: (2023)
Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus
por: Ananeva, Anastasia, et al.
Publicado: (2025)
por: Ananeva, Anastasia, et al.
Publicado: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
APCodec+: A Spectrum-Coding-Based High-Fidelity and High-Compression-Rate Neural Audio Codec with Staged Training Paradigm
por: Du, Hui-Peng, et al.
Publicado: (2024)
por: Du, Hui-Peng, et al.
Publicado: (2024)
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
por: Yoneyama, Reo, et al.
Publicado: (2025)
por: Yoneyama, Reo, et al.
Publicado: (2025)
Synthetic Data Domain Adaptation for ASR via LLM-based Text and Phonetic Respelling Augmentation
por: Yamashita, Natsuo, et al.
Publicado: (2026)
por: Yamashita, Natsuo, et al.
Publicado: (2026)
SGPA: Spectrogram-Guided Phonetic Alignment for Feasible Shapley Value Explanations in Multimodal Large Language Models
por: Pozorski, Paweł, et al.
Publicado: (2026)
por: Pozorski, Paweł, et al.
Publicado: (2026)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
por: Chu, Annie, et al.
Publicado: (2024)
por: Chu, Annie, et al.
Publicado: (2024)
Deep Audio Watermarks are Shallow: Limitations of Post-Hoc Watermarking Techniques for Speech
por: O'Reilly, Patrick, et al.
Publicado: (2025)
por: O'Reilly, Patrick, et al.
Publicado: (2025)
High-Fidelity Generative Audio Compression at 0.275kbps
por: Ma, Hao, et al.
Publicado: (2026)
por: Ma, Hao, et al.
Publicado: (2026)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
por: García, Hugo Flores, et al.
Publicado: (2024)
por: García, Hugo Flores, et al.
Publicado: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
por: Liu, Huadai, et al.
Publicado: (2024)
por: Liu, Huadai, et al.
Publicado: (2024)
InstructSing: High-Fidelity Singing Voice Generation via Instructing Yourself
por: Zeng, Chang, et al.
Publicado: (2024)
por: Zeng, Chang, et al.
Publicado: (2024)
STFTCodec: High-Fidelity Audio Compression through Time-Frequency Domain Representation
por: Feng, Tao, et al.
Publicado: (2025)
por: Feng, Tao, et al.
Publicado: (2025)
Towards High-Fidelity and Controllable Bioacoustic Generation via Enhanced Diffusion Learning
por: Song, Tianyu, et al.
Publicado: (2025)
por: Song, Tianyu, et al.
Publicado: (2025)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
por: Wang, Jin, et al.
Publicado: (2025)
por: Wang, Jin, et al.
Publicado: (2025)
High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
por: Lan, Gael Le, et al.
Publicado: (2024)
por: Lan, Gael Le, et al.
Publicado: (2024)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
por: Guimarães, Heitor R., et al.
Publicado: (2025)
por: Guimarães, Heitor R., et al.
Publicado: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
por: Du, Chenpeng, et al.
Publicado: (2022)
por: Du, Chenpeng, et al.
Publicado: (2022)
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
por: Guo, Yu-Ren, et al.
Publicado: (2025)
por: Guo, Yu-Ren, et al.
Publicado: (2025)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
por: Yuan, Jiajun, et al.
Publicado: (2025)
por: Yuan, Jiajun, et al.
Publicado: (2025)
A Generative-First Neural Audio Autoencoder
por: Casebeer, Jonah, et al.
Publicado: (2026)
por: Casebeer, Jonah, et al.
Publicado: (2026)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
por: Xin, Detai, et al.
Publicado: (2026)
por: Xin, Detai, et al.
Publicado: (2026)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
por: Chung, HaeChun
Publicado: (2025)
por: Chung, HaeChun
Publicado: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
por: Zhou, Kun, et al.
Publicado: (2024)
por: Zhou, Kun, et al.
Publicado: (2024)
SCRAPS: Speech Contrastive Representations of Acoustic and Phonetic Spaces
por: Vallés-Pérez, Ivan, et al.
Publicado: (2023)
por: Vallés-Pérez, Ivan, et al.
Publicado: (2023)
Ejemplares similares
-
Fine-Grained and Interpretable Neural Speech Editing
por: Morrison, Max, et al.
Publicado: (2024) -
Cross-domain Neural Pitch and Periodicity Estimation
por: Morrison, Max, et al.
Publicado: (2023) -
Combolutional Neural Networks
por: Churchwell, Cameron, et al.
Publicado: (2025) -
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
por: Li, Zirui, et al.
Publicado: (2025) -
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
por: Ahn, Sunghwan, et al.
Publicado: (2024)