A Real-Time Voice Activity Detection Based On Lightweight Neural
Fuente:
arXiv
Guardado en:
| Autores principales: | Jia, Jidong, Zhao, Pei, Wang, Di |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025)
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025)
Real-Time Pitch/F0 Detection Using Spectrogram Images and Convolutional Neural Networks
por: Zhao, Xufang, et al.
Publicado: (2025)
por: Zhao, Xufang, et al.
Publicado: (2025)
Selective Attention System (SAS): Device-Addressed Speech Detection for Real-Time On-Device Voice AI
por: Kim, David Joohun, et al.
Publicado: (2026)
por: Kim, David Joohun, et al.
Publicado: (2026)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
por: Wang, Jingyuan, et al.
Publicado: (2024)
por: Wang, Jingyuan, et al.
Publicado: (2024)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
por: Huang, Yubo, et al.
Publicado: (2024)
por: Huang, Yubo, et al.
Publicado: (2024)
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
por: Abrassart, Mathilde, et al.
Publicado: (2025)
por: Abrassart, Mathilde, et al.
Publicado: (2025)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
por: Anastassiou, Philip, et al.
Publicado: (2024)
por: Anastassiou, Philip, et al.
Publicado: (2024)
Audio-to-Image Encoding for Improved Voice Characteristic Detection Using Deep Convolutional Neural Networks
por: Atif, Youness
Publicado: (2025)
por: Atif, Youness
Publicado: (2025)
Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
por: Xu, Yiwen, et al.
Publicado: (2024)
por: Xu, Yiwen, et al.
Publicado: (2024)
VoiceWukong: Benchmarking Deepfake Voice Detection
por: Yan, Ziwei, et al.
Publicado: (2024)
por: Yan, Ziwei, et al.
Publicado: (2024)
SingFake: Singing Voice Deepfake Detection
por: Zang, Yongyi, et al.
Publicado: (2023)
por: Zang, Yongyi, et al.
Publicado: (2023)
Deepfake Detection of Singing Voices With Whisper Encodings
por: Sharma, Falguni, et al.
Publicado: (2025)
por: Sharma, Falguni, et al.
Publicado: (2025)
VoiceGRPO: Modern MoE Transformers with Group Relative Policy Optimization GRPO for AI Voice Health Care Applications on Voice Pathology Detection
por: Togootogtokh, Enkhtogtokh, et al.
Publicado: (2025)
por: Togootogtokh, Enkhtogtokh, et al.
Publicado: (2025)
Physics-Guided Deepfake Detection for Voice Authentication Systems
por: Mohammadi, Alireza, et al.
Publicado: (2025)
por: Mohammadi, Alireza, et al.
Publicado: (2025)
Voice Attribute Editing with Text Prompt
por: Sheng, Zhengyan, et al.
Publicado: (2024)
por: Sheng, Zhengyan, et al.
Publicado: (2024)
A Novel Labeled Human Voice Signal Dataset for Misbehavior Detection
por: Raza, Ali, et al.
Publicado: (2024)
por: Raza, Ali, et al.
Publicado: (2024)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
por: Zhao, Junchuan, et al.
Publicado: (2025)
por: Zhao, Junchuan, et al.
Publicado: (2025)
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
por: Sheng, Zhengyan, et al.
Publicado: (2025)
por: Sheng, Zhengyan, et al.
Publicado: (2025)
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation
por: Chen, Guo, et al.
Publicado: (2025)
por: Chen, Guo, et al.
Publicado: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
por: Guragain, Anmol, et al.
Publicado: (2024)
por: Guragain, Anmol, et al.
Publicado: (2024)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
por: Zheng, Junjie, et al.
Publicado: (2025)
por: Zheng, Junjie, et al.
Publicado: (2025)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
por: Li, Renyuan, et al.
Publicado: (2025)
por: Li, Renyuan, et al.
Publicado: (2025)
GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
por: Wu, Fan, et al.
Publicado: (2025)
por: Wu, Fan, et al.
Publicado: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
por: Xu, Rixi, et al.
Publicado: (2026)
por: Xu, Rixi, et al.
Publicado: (2026)
A New Approach to Voice Authenticity
por: Müller, Nicolas M., et al.
Publicado: (2024)
por: Müller, Nicolas M., et al.
Publicado: (2024)
HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
por: Nejad, Mahsa Ghazvini, et al.
Publicado: (2025)
por: Nejad, Mahsa Ghazvini, et al.
Publicado: (2025)
Voice Cloning: Comprehensive Survey
por: Azzuni, Hussam, et al.
Publicado: (2025)
por: Azzuni, Hussam, et al.
Publicado: (2025)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
por: Huang, Jiawei, et al.
Publicado: (2024)
por: Huang, Jiawei, et al.
Publicado: (2024)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
por: Yao, Wenhan, et al.
Publicado: (2024)
por: Yao, Wenhan, et al.
Publicado: (2024)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
por: Lu, Shenghui, et al.
Publicado: (2025)
por: Lu, Shenghui, et al.
Publicado: (2025)
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
por: Yao, Wenhan, et al.
Publicado: (2024)
por: Yao, Wenhan, et al.
Publicado: (2024)
Mitigating Latent Mismatch in cVAE-Based Singing Voice Synthesis via Flow Matching
por: Yun, Minhyeok, et al.
Publicado: (2026)
por: Yun, Minhyeok, et al.
Publicado: (2026)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
por: Zhang, Leying, et al.
Publicado: (2026)
por: Zhang, Leying, et al.
Publicado: (2026)
Asynchronous Voice Anonymization Using Adversarial Perturbation On Speaker Embedding
por: Wang, Rui, et al.
Publicado: (2024)
por: Wang, Rui, et al.
Publicado: (2024)
LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
por: Zou, Wenhao, et al.
Publicado: (2026)
por: Zou, Wenhao, et al.
Publicado: (2026)
Mozart's Touch: A Lightweight Multi-modal Music Generation Framework Based on Pre-Trained Large Models
por: Li, Jiajun, et al.
Publicado: (2024)
por: Li, Jiajun, et al.
Publicado: (2024)
LAPS-Diff: A Diffusion-Based Framework for Singing Voice Synthesis With Language Aware Prosody-Style Guided Learning
por: Dhar, Sandipan, et al.
Publicado: (2025)
por: Dhar, Sandipan, et al.
Publicado: (2025)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
por: Bai, Bingsong, et al.
Publicado: (2024)
por: Bai, Bingsong, et al.
Publicado: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
por: Pan, Yu, et al.
Publicado: (2024)
por: Pan, Yu, et al.
Publicado: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
por: Du, Zhihao, et al.
Publicado: (2025)
por: Du, Zhihao, et al.
Publicado: (2025)
Ejemplares similares
-
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
por: Chinchmalatpure, Prajwal, et al.
Publicado: (2025) -
Real-Time Pitch/F0 Detection Using Spectrogram Images and Convolutional Neural Networks
por: Zhao, Xufang, et al.
Publicado: (2025) -
Selective Attention System (SAS): Device-Addressed Speech Detection for Real-Time On-Device Voice AI
por: Kim, David Joohun, et al.
Publicado: (2026) -
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
por: Wang, Jingyuan, et al.
Publicado: (2024) -
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
por: Huang, Yubo, et al.
Publicado: (2024)