SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Mehta, Shivam, Liu, Yingru, Tang, Zhenyu, Peng, Kainan, Manohar, Vimal, Zhang, Shun, Seltzer, Mike, He, Qing, Ma, Mingbo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024)
by: Clapham, John, et al.
Published: (2024)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
by: Semenov, Andrei, et al.
Published: (2024)
by: Semenov, Andrei, et al.
Published: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
by: Mehta, Shivam, et al.
Published: (2025)
by: Mehta, Shivam, et al.
Published: (2025)
Uncovering Population PK Covariates from VAE-Generated Latent Spaces
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
Adaptive Riemannian Graph Neural Networks
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection
by: Wang, Xudong, et al.
Published: (2025)
by: Wang, Xudong, et al.
Published: (2025)
A Spatio-Temporal Deep Learning Approach For High-Resolution Gridded Monsoon Prediction
by: Borah, Parashjyoti, et al.
Published: (2026)
by: Borah, Parashjyoti, et al.
Published: (2026)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
by: Li, Pengcheng, et al.
Published: (2024)
by: Li, Pengcheng, et al.
Published: (2024)
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation
by: Perazzolo, Diego, et al.
Published: (2025)
by: Perazzolo, Diego, et al.
Published: (2025)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
by: Sharma, Anika, et al.
Published: (2025)
by: Sharma, Anika, et al.
Published: (2025)
Regularisation in neural networks: a survey and empirical analysis of approaches
by: Opperman, Christiaan P., et al.
Published: (2026)
by: Opperman, Christiaan P., et al.
Published: (2026)
Alternative Local Discriminant Bases Using Empirical Expectation and Variance Estimation
by: Fossgaard, Eirik
Published: (1999)
by: Fossgaard, Eirik
Published: (1999)
FedAlign: Federated Domain Generalization with Cross-Client Feature Alignment
by: Gupta, Sunny, et al.
Published: (2025)
by: Gupta, Sunny, et al.
Published: (2025)
Expanding continual few-shot learning benchmarks to include recognition of specific instances
by: Kowadlo, Gideon, et al.
Published: (2022)
by: Kowadlo, Gideon, et al.
Published: (2022)
Balancing the Scales: A Comprehensive Study on Tackling Class Imbalance in Binary Classification
by: Abdelhamid, Mohamed, et al.
Published: (2024)
by: Abdelhamid, Mohamed, et al.
Published: (2024)
Discriminative Subspace Emersion from learning feature relevances across different populations
by: Canducci, Marco, et al.
Published: (2025)
by: Canducci, Marco, et al.
Published: (2025)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
by: Bouzid, Hamza, et al.
Published: (2023)
by: Bouzid, Hamza, et al.
Published: (2023)
ScaleMAP: Preserving Local Density and Neighborhood Structure in Low-Dimensional Embeddings
by: Poorna, Rajas, et al.
Published: (2026)
by: Poorna, Rajas, et al.
Published: (2026)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
by: Lapin, Mykyta, et al.
Published: (2025)
by: Lapin, Mykyta, et al.
Published: (2025)
CASE: Contrastive Activation for Saliency Estimation
by: Williamson, Dane, et al.
Published: (2025)
by: Williamson, Dane, et al.
Published: (2025)
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
by: Mehta, Shivam, et al.
Published: (2024)
by: Mehta, Shivam, et al.
Published: (2024)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
by: Du, Guanchen, et al.
Published: (2025)
by: Du, Guanchen, et al.
Published: (2025)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
by: Mehta, Diksha, et al.
Published: (2025)
by: Mehta, Diksha, et al.
Published: (2025)
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
by: Anand, Emile, et al.
Published: (2026)
by: Anand, Emile, et al.
Published: (2026)
PRACH Preamble Detection as a Multi-Class Classification Problem: A Machine Learning Approach Using SVM
by: Ferenc, F., et al.
Published: (2025)
by: Ferenc, F., et al.
Published: (2025)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
by: Martin, Michael R., et al.
Published: (2025)
by: Martin, Michael R., et al.
Published: (2025)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
by: Banthia, Saumya, et al.
Published: (2020)
by: Banthia, Saumya, et al.
Published: (2020)
Precision at Scale: Domain-Specific Datasets On-Demand
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
by: Rodríguez-de-Vera, Jesús M, et al.
Published: (2024)
Prior-Aligned Data Cleaning for Tabular Foundation Models
by: Berti-Equille, Laure
Published: (2026)
by: Berti-Equille, Laure
Published: (2026)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
by: Kalušev, Vladimir, et al.
Published: (2025)
by: Kalušev, Vladimir, et al.
Published: (2025)
Matcha-TTS: A fast TTS architecture with conditional flow matching
by: Mehta, Shivam, et al.
Published: (2023)
by: Mehta, Shivam, et al.
Published: (2023)
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
by: Shaw, Ankit Kumar, et al.
Published: (2025)
by: Shaw, Ankit Kumar, et al.
Published: (2025)
DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology
by: Dimasaka, Joshua, et al.
Published: (2025)
by: Dimasaka, Joshua, et al.
Published: (2025)
VulScribeR: Exploring RAG-based Vulnerability Augmentation with LLMs
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
by: Daneshvar, Seyed Shayan, et al.
Published: (2024)
In-situ Autoguidance: Eliciting Self-Correction in Diffusion Models
by: Gu, Enhao, et al.
Published: (2025)
by: Gu, Enhao, et al.
Published: (2025)
Impulsive pattern recognition of a myoelectric hand via Dynamic Time Warping
by: Kadilar, Mustafa Can, et al.
Published: (2025)
by: Kadilar, Mustafa Can, et al.
Published: (2025)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
by: Chien, Sheng-You, et al.
Published: (2026)
by: Chien, Sheng-You, et al.
Published: (2026)
Pairwise Spatiotemporal Partial Trajectory Matching for Co-movement Analysis
by: Cardei, Maria, et al.
Published: (2024)
by: Cardei, Maria, et al.
Published: (2024)
Plug In and Learn: Federated Intelligence over a Smart Grid of Models
by: Abdurakhmanova, S., et al.
Published: (2023)
by: Abdurakhmanova, S., et al.
Published: (2023)
Similar Items
-
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024) -
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
by: Semenov, Andrei, et al.
Published: (2024) -
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
by: Mehta, Shivam, et al.
Published: (2025) -
Uncovering Population PK Covariates from VAE-Generated Latent Spaces
by: Perazzolo, Diego, et al.
Published: (2025) -
Adaptive Riemannian Graph Neural Networks
by: Wang, Xudong, et al.
Published: (2025)