Text-Aware Adapter for Few-Shot Keyword Spotting
Fuente:
arXiv
Saved in:
| Main Authors: | Jung, Youngmoon, Lee, Jinyoung, Lee, Seungjin, Jung, Myunghun, Lee, Yong-Hyeok, Cho, Hoon-Young |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relational Proxy Loss for Audio-Text based Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2024)
by: Jung, Youngmoon, et al.
Published: (2024)
MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2026)
by: Jung, Youngmoon, et al.
Published: (2026)
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2025)
by: Jung, Youngmoon, et al.
Published: (2025)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
by: Jin, Sichen, et al.
Published: (2024)
by: Jin, Sichen, et al.
Published: (2024)
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
by: Cho, Hyunjae, et al.
Published: (2024)
by: Cho, Hyunjae, et al.
Published: (2024)
DAME: Duration-Aware Matryoshka Embedding for Duration-Robust Speaker Verification
by: Jung, Youngmoon, et al.
Published: (2026)
by: Jung, Youngmoon, et al.
Published: (2026)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
by: Bae, Hanbin, et al.
Published: (2024)
by: Bae, Hanbin, et al.
Published: (2024)
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders
by: Kim, Seungbae, et al.
Published: (2025)
by: Kim, Seungbae, et al.
Published: (2025)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
by: Lee, Dongheon, et al.
Published: (2023)
by: Lee, Dongheon, et al.
Published: (2023)
PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
by: Lee, Sang-Hoon, et al.
Published: (2024)
by: Lee, Sang-Hoon, et al.
Published: (2024)
Accelerating High-Fidelity Waveform Generation via Adversarial Flow Matching Optimization
by: Lee, Sang-Hoon, et al.
Published: (2024)
by: Lee, Sang-Hoon, et al.
Published: (2024)
Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
by: Wu, Haibin, et al.
Published: (2024)
by: Wu, Haibin, et al.
Published: (2024)
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
by: Kuzmin, Nikita, et al.
Published: (2026)
by: Kuzmin, Nikita, et al.
Published: (2026)
When Humans Growl and Birds Speak: High-Fidelity Voice Conversion from Human to Animal and Designed Sounds
by: Kang, Minsu, et al.
Published: (2025)
by: Kang, Minsu, et al.
Published: (2025)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
by: Nishu, Kumari, et al.
Published: (2024)
by: Nishu, Kumari, et al.
Published: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
by: Lee, Jin Woo, et al.
Published: (2024)
by: Lee, Jin Woo, et al.
Published: (2024)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
by: Kim, Jaeyeon, et al.
Published: (2024)
by: Kim, Jaeyeon, et al.
Published: (2024)
An Investigation of Noise Robustness for Flow-Matching-Based Zero-Shot TTS
by: Wang, Xiaofei, et al.
Published: (2024)
by: Wang, Xiaofei, et al.
Published: (2024)
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
by: Kuzmin, Nikita, et al.
Published: (2026)
by: Kuzmin, Nikita, et al.
Published: (2026)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
by: Choi, Woongjib, et al.
Published: (2025)
by: Choi, Woongjib, et al.
Published: (2025)
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment
by: Cai, Huanchen, et al.
Published: (2026)
by: Cai, Huanchen, et al.
Published: (2026)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
by: Lee, Jihwan, et al.
Published: (2024)
by: Lee, Jihwan, et al.
Published: (2024)
DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment
by: Oh, Hyung-Seok, et al.
Published: (2024)
by: Oh, Hyung-Seok, et al.
Published: (2024)
EchoScan: Scanning Complex Room Geometries via Acoustic Echoes
by: Yeon, Inmo, et al.
Published: (2023)
by: Yeon, Inmo, et al.
Published: (2023)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
by: Huang, Kuan-Tang, et al.
Published: (2026)
by: Huang, Kuan-Tang, et al.
Published: (2026)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
by: Choi, Ha-Yeong, et al.
Published: (2025)
by: Choi, Ha-Yeong, et al.
Published: (2025)
Triage knowledge distillation for speaker verification
by: Kim, Ju-ho, et al.
Published: (2026)
by: Kim, Ju-ho, et al.
Published: (2026)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
by: Kim, Ji-Hoon, et al.
Published: (2024)
by: Kim, Ji-Hoon, et al.
Published: (2024)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
by: Lee, Jihwan, et al.
Published: (2025)
by: Lee, Jihwan, et al.
Published: (2025)
A Knowledge-Driven Approach to Music Segmentation, Music Source Separation and Cinematic Audio Source Separation
by: Ho, Chun-wei, et al.
Published: (2026)
by: Ho, Chun-wei, et al.
Published: (2026)
Self-supervised Learning for Acoustic Few-Shot Classification
by: Liang, Jingyong, et al.
Published: (2024)
by: Liang, Jingyong, et al.
Published: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
by: Cho, Deok-Hyeon, et al.
Published: (2024)
by: Cho, Deok-Hyeon, et al.
Published: (2024)
T-FOLEY: A Controllable Waveform-Domain Diffusion Model for Temporal-Event-Guided Foley Sound Synthesis
by: Chung, Yoonjin, et al.
Published: (2024)
by: Chung, Yoonjin, et al.
Published: (2024)
Synth4Kws: Synthesized Speech for User Defined Keyword Spotting in Low Resource Environments
by: Zhu, Pai, et al.
Published: (2024)
by: Zhu, Pai, et al.
Published: (2024)
Advanced Signal Analysis in Detecting Replay Attacks for Automatic Speaker Verification Systems
by: Kuang, Lee Shih
Published: (2024)
by: Kuang, Lee Shih
Published: (2024)
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
by: Jin, Hojun, et al.
Published: (2025)
by: Jin, Hojun, et al.
Published: (2025)
Probing Cross-modal Information Hubs in Audio-Visual LLMs
by: Jung, Jihoo, et al.
Published: (2026)
by: Jung, Jihoo, et al.
Published: (2026)
Transferable Adversarial Attacks against ASR
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
Crowdsourced Multilingual Speech Intelligibility Testing
by: Lechler, Laura, et al.
Published: (2024)
by: Lechler, Laura, et al.
Published: (2024)
Similar Items
-
Relational Proxy Loss for Audio-Text based Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2024) -
MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2026) -
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
by: Jung, Youngmoon, et al.
Published: (2025) -
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
by: Jin, Sichen, et al.
Published: (2024) -
JenGAN: Stacked Shifted Filters in GAN-Based Speech Synthesis
by: Cho, Hyunjae, et al.
Published: (2024)