CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
Fuente:
arXiv
Salvato in:
| Autori principali: | Jin, Sichen, Jung, Youngmoon, Lee, Seungjin, Roh, Jaeyoung, Han, Changwoo, Cho, Hoonyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2026)
di: Jung, Youngmoon, et al.
Pubblicazione: (2026)
Relational Proxy Loss for Audio-Text based Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2025)
di: Jung, Youngmoon, et al.
Pubblicazione: (2025)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
di: Xi, Yu, et al.
Pubblicazione: (2024)
di: Xi, Yu, et al.
Pubblicazione: (2024)
Text-Aware Adapter for Few-Shot Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)
Streaming Keyword Spotting Boosted by Cross-layer Discrimination Consistency
di: Xi, Yu, et al.
Pubblicazione: (2024)
di: Xi, Yu, et al.
Pubblicazione: (2024)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
di: Xi, Yu, et al.
Pubblicazione: (2024)
di: Xi, Yu, et al.
Pubblicazione: (2024)
Keyword Mamba: Spoken Keyword Spotting with State Space Models
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
di: Ding, Hanyu, et al.
Pubblicazione: (2025)
Effective Integration of KAN for Keyword Spotting
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
Multichannel Keyword Spotting for Noisy Conditions
di: Saladukha, Dzmitry, et al.
Pubblicazione: (2025)
di: Saladukha, Dzmitry, et al.
Pubblicazione: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
di: Nakagome, Yu, et al.
Pubblicazione: (2025)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
di: Liu, Hanwen, et al.
Pubblicazione: (2026)
di: Liu, Hanwen, et al.
Pubblicazione: (2026)
Self-Learning for Personalized Keyword Spotting on Ultra-Low-Power Audio Sensors
di: Rusci, Manuele, et al.
Pubblicazione: (2024)
di: Rusci, Manuele, et al.
Pubblicazione: (2024)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaeyeon, et al.
Pubblicazione: (2024)
Effective User-defined Keyword Spotting with Dual-stage Matching, Multi-modal Enrollment, and Continual Adaptation
di: Ai, Zhiqi, et al.
Pubblicazione: (2026)
di: Ai, Zhiqi, et al.
Pubblicazione: (2026)
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding
di: Xi, Yu, et al.
Pubblicazione: (2025)
di: Xi, Yu, et al.
Pubblicazione: (2025)
TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
di: Xi, Yu, et al.
Pubblicazione: (2024)
di: Xi, Yu, et al.
Pubblicazione: (2024)
Phoneme-Level Contrastive Learning for User-Defined Keyword Spotting with Flexible Enrollment
di: Kewei, Li, et al.
Pubblicazione: (2024)
di: Kewei, Li, et al.
Pubblicazione: (2024)
Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
di: Xi, Yu, et al.
Pubblicazione: (2025)
di: Xi, Yu, et al.
Pubblicazione: (2025)
Sparse Binarization for Fast Keyword Spotting
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
di: Li, Longhao, et al.
Pubblicazione: (2025)
di: Li, Longhao, et al.
Pubblicazione: (2025)
On-Device Domain Learning for Keyword Spotting on Low-Power Extreme Edge Embedded Systems
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
di: Wu, Shih-Lun, et al.
Pubblicazione: (2023)
di: Wu, Shih-Lun, et al.
Pubblicazione: (2023)
Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
di: Garai, Soumen, et al.
Pubblicazione: (2025)
di: Garai, Soumen, et al.
Pubblicazione: (2025)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
di: Zhou, Jiaming, et al.
Pubblicazione: (2023)
Does Single-channel Speech Enhancement Improve Keyword Spotting Accuracy? A Case Study
di: Brueggeman, Avamarie, et al.
Pubblicazione: (2023)
di: Brueggeman, Avamarie, et al.
Pubblicazione: (2023)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
CR-CTC: Consistency regularization on CTC for improved speech recognition
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
di: Yao, Zengwei, et al.
Pubblicazione: (2024)
Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio Effects
di: Chu, Annie, et al.
Pubblicazione: (2024)
di: Chu, Annie, et al.
Pubblicazione: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
di: Hou, Junfeng, et al.
Pubblicazione: (2024)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
EdgeSpot: Efficient and High-Performance Few-Shot Model for Keyword Spotting
di: Buyuksolak, Oguzhan, et al.
Pubblicazione: (2026)
di: Buyuksolak, Oguzhan, et al.
Pubblicazione: (2026)
Multiple-Instance, Cascaded Classification for Keyword Spotting in Narrow-Band Audio
di: AbdulKader, Ahmad, et al.
Pubblicazione: (2017)
di: AbdulKader, Ahmad, et al.
Pubblicazione: (2017)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
di: Nishu, Kumari, et al.
Pubblicazione: (2024)
di: Nishu, Kumari, et al.
Pubblicazione: (2024)
LLM-Synth4KWS: Scalable Automatic Generation and Synthesis of Confusable Data for Custom Keyword Spotting
di: Zhu, Pai, et al.
Pubblicazione: (2025)
di: Zhu, Pai, et al.
Pubblicazione: (2025)
Quantization-Based Score Calibration for Few-Shot Keyword Spotting with Dynamic Time Warping in Noisy Environments
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2025)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2025)
Streaming Audio Transformers for Online Audio Tagging
di: Dinkel, Heinrich, et al.
Pubblicazione: (2023)
di: Dinkel, Heinrich, et al.
Pubblicazione: (2023)
Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
di: Park, Hyun Jin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
MATE: Matryoshka Audio-Text Embeddings for Open-Vocabulary Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2026) -
Relational Proxy Loss for Audio-Text based Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024) -
Adversarial Deep Metric Learning for Cross-Modal Audio-Text Alignment in Open-Vocabulary Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2025) -
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
di: Xi, Yu, et al.
Pubblicazione: (2024) -
Text-Aware Adapter for Few-Shot Keyword Spotting
di: Jung, Youngmoon, et al.
Pubblicazione: (2024)