Saved in:
| Main Authors: | Kang, Taein, Han, Soyul, Choi, Sunmook, Seo, Jaejin, Chung, Sanghyeok, Lee, Seungeun, Oh, Seungsang, Kwak, Il-Youp |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.17127 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
Automatic classification of stop realisation with wav2vec2.0
by: Tanner, James, et al.
Published: (2025)
by: Tanner, James, et al.
Published: (2025)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
by: Bayerl, Sebastian P., et al.
Published: (2022)
by: Bayerl, Sebastian P., et al.
Published: (2022)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
by: Wang, Zhiyong, et al.
Published: (2024)
by: Wang, Zhiyong, et al.
Published: (2024)
A Comprehensive Corpus of Biomechanically Constrained Piano Chords: Generation, Analysis, and Implications for Voicing and Psychoacoustics
by: Ramani, Mahesh
Published: (2026)
by: Ramani, Mahesh
Published: (2026)
Fully Few-shot Class-incremental Audio Classification Using Multi-level Embedding Extractor and Ridge Regression Classifier
by: Si, Yongjie, et al.
Published: (2025)
by: Si, Yongjie, et al.
Published: (2025)
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
by: Zhu, Qiushi, et al.
Published: (2024)
by: Zhu, Qiushi, et al.
Published: (2024)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
by: Huo, Robin, et al.
Published: (2025)
by: Huo, Robin, et al.
Published: (2025)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
by: Li, Feng, et al.
Published: (2024)
by: Li, Feng, et al.
Published: (2024)
BEAT2AASIST model with layer fusion for ESDD 2026 Challenge
by: Chung, Sanghyeok, et al.
Published: (2025)
by: Chung, Sanghyeok, et al.
Published: (2025)
Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
by: Demir, Kubilay Can, et al.
Published: (2024)
by: Demir, Kubilay Can, et al.
Published: (2024)
Rhythmic segment analysis: Conceptualizing, visualizing, and measuring rhythmic data
by: Cornelissen, Bas
Published: (2026)
by: Cornelissen, Bas
Published: (2026)
Golden Tonnetz
by: Imai, Yusuke
Published: (2025)
by: Imai, Yusuke
Published: (2025)
An introduction to pitch strength in contemporary popular music analysis and production
by: Deruty, Emmanuel
Published: (2025)
by: Deruty, Emmanuel
Published: (2025)
Proliferating series by Jean Barraqué: a study and classification in mathematical terms
by: Tardón, Isabel, et al.
Published: (2026)
by: Tardón, Isabel, et al.
Published: (2026)
From Imitation to Innovation: The Divergent Paths of Techno in Germany and the USA
by: Ziemer, Tim, et al.
Published: (2025)
by: Ziemer, Tim, et al.
Published: (2025)
Methods for pitch analysis in contemporary popular music: Vitalic's use of tones that do not operate on the principle of acoustic resonance
by: Deruty, Emmanuel, et al.
Published: (2025)
by: Deruty, Emmanuel, et al.
Published: (2025)
wav2pos: Sound Source Localization using Masked Autoencoders
by: Berg, Axel, et al.
Published: (2024)
by: Berg, Axel, et al.
Published: (2024)
Detecting Spoof Voices in Asian Non-Native Speech: An Indonesian and Thai Case Study
by: Adila, Aulia, et al.
Published: (2024)
by: Adila, Aulia, et al.
Published: (2024)
UNMIXX: Untangling Highly Correlated Singing Voices Mixtures
by: Jung, Jihoo, et al.
Published: (2026)
by: Jung, Jihoo, et al.
Published: (2026)
An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
by: Zhang, You, et al.
Published: (2021)
by: Zhang, You, et al.
Published: (2021)
Rapidly Adapting to New Voice Spoofing: Few-Shot Detection of Synthesized Speech Under Distribution Shifts
by: Garg, Ashi, et al.
Published: (2025)
by: Garg, Ashi, et al.
Published: (2025)
voc2vec: A Foundation Model for Non-Verbal Vocalization
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
by: Kurnaz, Oğuzhan, et al.
Published: (2025)
by: Kurnaz, Oğuzhan, et al.
Published: (2025)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
by: Kwak, Doyeop, et al.
Published: (2026)
by: Kwak, Doyeop, et al.
Published: (2026)
LP-CFM: Perceptual Invariance-Aware Conditional Flow Matching for Speech Modeling
by: Kwak, Doyeop, et al.
Published: (2025)
by: Kwak, Doyeop, et al.
Published: (2025)
Improving Short Utterance Anti-Spoofing with AASIST2
by: Zhang, Yuxiang, et al.
Published: (2023)
by: Zhang, Yuxiang, et al.
Published: (2023)
Algebraic Structures in Microtonal Music
by: Flynn, Veronica, et al.
Published: (2025)
by: Flynn, Veronica, et al.
Published: (2025)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
by: Shao, Qijie, et al.
Published: (2025)
by: Shao, Qijie, et al.
Published: (2025)
VoiceCraft-Dub: Automated Video Dubbing with Neural Codec Language Models
by: Sung-Bin, Kim, et al.
Published: (2025)
by: Sung-Bin, Kim, et al.
Published: (2025)
Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
by: Zhang, Lin, et al.
Published: (2024)
by: Zhang, Lin, et al.
Published: (2024)
VoxSim: A perceptual voice similarity dataset
by: Ahn, Junseok, et al.
Published: (2024)
by: Ahn, Junseok, et al.
Published: (2024)
BUT Systems for WildSpoof Challenge: SASV in the Wild
by: Peng, Junyi, et al.
Published: (2025)
by: Peng, Junyi, et al.
Published: (2025)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
by: Attia, Ahmed Adel, et al.
Published: (2024)
by: Attia, Ahmed Adel, et al.
Published: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
by: Attia, Ahmed Adel, et al.
Published: (2024)
by: Attia, Ahmed Adel, et al.
Published: (2024)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
Optimizing a-DCF for Spoofing-Robust Speaker Verification
by: Kurnaz, Oğuzhan, et al.
Published: (2024)
by: Kurnaz, Oğuzhan, et al.
Published: (2024)
Token-based Attractors and Cross-attention in Spoof Diarization
by: Koo, Kyo-Won, et al.
Published: (2025)
by: Koo, Kyo-Won, et al.
Published: (2025)
When Spoof Detectors Travel: Evaluation Across 66 Languages in the Low-Resource Language Spoofing Corpus
by: Borodin, Kirill, et al.
Published: (2026)
by: Borodin, Kirill, et al.
Published: (2026)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
by: Abdullah, Abdulhady Abas, et al.
Published: (2025)
Similar Items
-
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
by: Guo, Yiwei, et al.
Published: (2024) -
Automatic classification of stop realisation with wav2vec2.0
by: Tanner, James, et al.
Published: (2025) -
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
by: Bayerl, Sebastian P., et al.
Published: (2022) -
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
by: Wang, Zhiyong, et al.
Published: (2024) -
A Comprehensive Corpus of Biomechanically Constrained Piano Chords: Generation, Analysis, and Implications for Voicing and Psychoacoustics
by: Ramani, Mahesh
Published: (2026)