VoxSim: A perceptual voice similarity dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Ahn, Junseok, Kim, Youkyum, Choi, Yeunju, Kwak, Doyeop, Kim, Ji-Hoon, Mun, Seongkyu, Chung, Joon Son |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaptVC: High Quality Voice Conversion with Adaptive Learning
by: Kim, Jaehun, et al.
Published: (2025)
by: Kim, Jaehun, et al.
Published: (2025)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
by: Kim, Ji-Hoon, et al.
Published: (2025)
by: Kim, Ji-Hoon, et al.
Published: (2025)
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
by: Kwak, Doyeop, et al.
Published: (2025)
by: Kwak, Doyeop, et al.
Published: (2025)
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
by: Kwak, Doyeop, et al.
Published: (2025)
by: Kwak, Doyeop, et al.
Published: (2025)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
by: Jung, Jihoo, et al.
Published: (2026)
by: Jung, Jihoo, et al.
Published: (2026)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
Bridging the Gap between Audio and Text using Parallel-attention for User-defined Keyword Spotting
by: Kim, Youkyum, et al.
Published: (2024)
by: Kim, Youkyum, et al.
Published: (2024)
InfiniteAudio: Infinite-Length Audio Generation with Consistency
by: Jung, Chaeyoung, et al.
Published: (2025)
by: Jung, Chaeyoung, et al.
Published: (2025)
Faces that Speak: Jointly Synthesising Talking Face and Speech from Text
by: Jang, Youngjoon, et al.
Published: (2024)
by: Jang, Youngjoon, et al.
Published: (2024)
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
by: Nguyen, Tan Dat, et al.
Published: (2026)
by: Nguyen, Tan Dat, et al.
Published: (2026)
FlowAVSE: Efficient Audio-Visual Speech Enhancement with Conditional Flow Matching
by: Jung, Chaeyoung, et al.
Published: (2024)
by: Jung, Chaeyoung, et al.
Published: (2024)
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
by: Nguyen, Tan Dat, et al.
Published: (2025)
by: Nguyen, Tan Dat, et al.
Published: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
Probing Cross-modal Information Hubs in Audio-Visual LLMs
by: Jung, Jihoo, et al.
Published: (2026)
by: Jung, Jihoo, et al.
Published: (2026)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
by: Kim, Ji-Hoon, et al.
Published: (2023)
by: Kim, Ji-Hoon, et al.
Published: (2023)
Latent Filling: Latent Space Data Augmentation for Zero-shot Speech Synthesis
by: Bae, Jae-Sung, et al.
Published: (2023)
by: Bae, Jae-Sung, et al.
Published: (2023)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
by: Jung, Jaemin, et al.
Published: (2024)
by: Jung, Jaemin, et al.
Published: (2024)
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
by: Rautenberg, Frederik, et al.
Published: (2025)
by: Rautenberg, Frederik, et al.
Published: (2025)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
SCORE: Scaling audio generation using Standardized COmposite REwards
by: Jung, Jaemin, et al.
Published: (2025)
by: Jung, Jaemin, et al.
Published: (2025)
V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow
by: Choi, Jeongsoo, et al.
Published: (2024)
by: Choi, Jeongsoo, et al.
Published: (2024)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech
by: Kim, Ji-Hoon, et al.
Published: (2025)
by: Kim, Ji-Hoon, et al.
Published: (2025)
The VoxCeleb Speaker Recognition Challenge: A Retrospective
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
Evaluating voice anonymisation using similarity rank disclosure
by: Chandra, Shilpa, et al.
Published: (2026)
by: Chandra, Shilpa, et al.
Published: (2026)
Investigation of perceptual music similarity focusing on each instrumental part
by: Hashizume, Yuka, et al.
Published: (2025)
by: Hashizume, Yuka, et al.
Published: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
by: Lee, Jaesong, et al.
Published: (2024)
by: Lee, Jaesong, et al.
Published: (2024)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
Team HYU ASML ROBOVOX SP Cup 2024 System Description
by: Choi, Jeong-Hwan, et al.
Published: (2024)
by: Choi, Jeong-Hwan, et al.
Published: (2024)
HeightCeleb - an enrichment of VoxCeleb dataset with speaker height information
by: Kacprzak, Stanisław, et al.
Published: (2024)
by: Kacprzak, Stanisław, et al.
Published: (2024)
Speech Corpus for Korean Children with Autism Spectrum Disorder: Towards Automatic Assessment Systems
by: Lee, Seonwoo, et al.
Published: (2024)
by: Lee, Seonwoo, et al.
Published: (2024)
Curriculum learning for self-supervised speaker verification
by: Heo, Hee-Soo, et al.
Published: (2022)
by: Heo, Hee-Soo, et al.
Published: (2022)
When Vision Models Meet Parameter Efficient Look-Aside Adapters Without Large-Scale Audio Pretraining
by: Yeo, Juan, et al.
Published: (2024)
by: Yeo, Juan, et al.
Published: (2024)
VoxATtack: A Multimodal Attack on Voice Anonymization Systems
by: Aloradi, Ahmad, et al.
Published: (2025)
by: Aloradi, Ahmad, et al.
Published: (2025)
Triage knowledge distillation for speaker verification
by: Kim, Ju-ho, et al.
Published: (2026)
by: Kim, Ju-ho, et al.
Published: (2026)
Mamba2 Meets Silence: Robust Vocal Source Separation for Sparse Regions
by: Kim, Euiyeon, et al.
Published: (2025)
by: Kim, Euiyeon, et al.
Published: (2025)
VoxEffects: A Speech-Oriented Audio Effects Dataset and Benchmark
by: Zhang, Zhe, et al.
Published: (2026)
by: Zhang, Zhe, et al.
Published: (2026)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
by: Lee, Joun Yeop, et al.
Published: (2024)
by: Lee, Joun Yeop, et al.
Published: (2024)
Similar Items
-
AdaptVC: High Quality Voice Conversion with Adaptive Learning
by: Kim, Jaehun, et al.
Published: (2025) -
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026) -
TAVID: Text-Driven Audio-Visual Interactive Dialogue Generation
by: Kim, Ji-Hoon, et al.
Published: (2025) -
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
by: Kwak, Doyeop, et al.
Published: (2025) -
EDNet: A Versatile Speech Enhancement Framework with Gating Mamba Mechanism and Phase Shift-Invariant Training
by: Kwak, Doyeop, et al.
Published: (2025)