Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Tianchi, Truong, Duc-Tuan, Das, Rohan Kumar, Lee, Kong Aik, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
par: Xiao, Yang, et autres
Publié: (2025)
par: Xiao, Yang, et autres
Publié: (2025)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
par: Li, Junjie, et autres
Publié: (2025)
par: Li, Junjie, et autres
Publié: (2025)
Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification
par: Liu, Tianchi, et autres
Publié: (2023)
par: Liu, Tianchi, et autres
Publié: (2023)
Room Impulse Responses help attackers to evade Deep Fake Detection
par: Luong, Hieu-Thi, et autres
Publié: (2024)
par: Luong, Hieu-Thi, et autres
Publié: (2024)
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
par: Truong, Duc-Tuan, et autres
Publié: (2023)
par: Truong, Duc-Tuan, et autres
Publié: (2023)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
par: Truong, Duc-Tuan, et autres
Publié: (2024)
par: Truong, Duc-Tuan, et autres
Publié: (2024)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
par: Trachu, Thanapat, et autres
Publié: (2025)
par: Trachu, Thanapat, et autres
Publié: (2025)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
par: Pan, Zihan, et autres
Publié: (2024)
par: Pan, Zihan, et autres
Publié: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
par: Ma, Yi, et autres
Publié: (2026)
par: Ma, Yi, et autres
Publié: (2026)
Gradient weighting for speaker verification in extremely low Signal-to-Noise Ratio
par: Ma, Yi, et autres
Publié: (2024)
par: Ma, Yi, et autres
Publié: (2024)
Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
par: Wang, Shuai, et autres
Publié: (2024)
par: Wang, Shuai, et autres
Publié: (2024)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
par: Luong, Hieu-Thi, et autres
Publié: (2024)
par: Luong, Hieu-Thi, et autres
Publié: (2024)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
par: Liu, Tianchi, et autres
Publié: (2024)
par: Liu, Tianchi, et autres
Publié: (2024)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
par: Li, Junjie, et autres
Publié: (2024)
par: Li, Junjie, et autres
Publié: (2024)
How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
par: Liu, Tianchi, et autres
Publié: (2024)
par: Liu, Tianchi, et autres
Publié: (2024)
Robust Localization of Partially Fake Speech: Metrics and Out-of-Domain Evaluation
par: Luong, Hieu-Thi, et autres
Publié: (2025)
par: Luong, Hieu-Thi, et autres
Publié: (2025)
Adversarial speech for voice privacy protection from Personalized Speech generation
par: Chen, Shihao, et autres
Publié: (2024)
par: Chen, Shihao, et autres
Publié: (2024)
Multi-modal Speech Enhancement with Limited Electromyography Channels
par: Feng, Fuyuan, et autres
Publié: (2025)
par: Feng, Fuyuan, et autres
Publié: (2025)
Can Emotion Fool Anti-spoofing?
par: Mahapatra, Aurosweta, et autres
Publié: (2025)
par: Mahapatra, Aurosweta, et autres
Publié: (2025)
Study of Lightweight Transformer Architectures for Single-Channel Speech Enhancement
par: Zhao, Haixin, et autres
Publié: (2025)
par: Zhao, Haixin, et autres
Publié: (2025)
CompSpoof: A Dataset and Joint Learning Framework for Component-Level Audio Anti-spoofing Countermeasures
par: Zhang, Xueping, et autres
Publié: (2025)
par: Zhang, Xueping, et autres
Publié: (2025)
Cosine Scoring with Uncertainty for Neural Speaker Embedding
par: Wang, Qiongqiong, et autres
Publié: (2024)
par: Wang, Qiongqiong, et autres
Publié: (2024)
ChunkFormer: Masked Chunking Conformer For Long-Form Speech Transcription
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
Enhancing Anti-spoofing Countermeasures Robustness through Joint Optimization and Transfer Learning
par: Wang, Yikang, et autres
Publié: (2024)
par: Wang, Yikang, et autres
Publié: (2024)
Device Feature based on Graph Fourier Transformation with Logarithmic Processing For Detection of Replay Speech Attacks
par: He, Mingrui, et autres
Publié: (2024)
par: He, Mingrui, et autres
Publié: (2024)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
Listen, Analyze, and Adapt to Learn New Attacks: An Exemplar-Free Class Incremental Learning Method for Audio Deepfake Source Tracing
par: Xiao, Yang, et autres
Publié: (2025)
par: Xiao, Yang, et autres
Publié: (2025)
Where's That Voice Coming? Continual Learning for Sound Source Localization
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
TF-Mamba: A Time-Frequency Network for Sound Source Localization
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
par: Xiao, Yang, et autres
Publié: (2024)
par: Xiao, Yang, et autres
Publié: (2024)
Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
par: Wang, Xin, et autres
Publié: (2024)
par: Wang, Xin, et autres
Publié: (2024)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
par: Lin, Weiwei, et autres
Publié: (2024)
par: Lin, Weiwei, et autres
Publié: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
par: Inoue, Sho, et autres
Publié: (2024)
par: Inoue, Sho, et autres
Publié: (2024)
ExPO: Explainable Phonetic Trait-Oriented Network for Speaker Verification
par: Ma, Yi, et autres
Publié: (2025)
par: Ma, Yi, et autres
Publié: (2025)
VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
par: Kim, Heeseung, et autres
Publié: (2024)
par: Kim, Heeseung, et autres
Publié: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
par: Inoue, Sho, et autres
Publié: (2025)
par: Inoue, Sho, et autres
Publié: (2025)
Documents similaires
-
RawTFNet: A Lightweight CNN Architecture for Speech Anti-spoofing
par: Xiao, Yang, et autres
Publié: (2025) -
Xi+: Uncertainty Supervision for Robust Speaker Embedding
par: Li, Junjie, et autres
Publié: (2025) -
Golden Gemini is All You Need: Finding the Sweet Spots for Speaker Verification
par: Liu, Tianchi, et autres
Publié: (2023) -
Room Impulse Responses help attackers to evade Deep Fake Detection
par: Luong, Hieu-Thi, et autres
Publié: (2024) -
Emphasized Non-Target Speaker Knowledge in Knowledge Distillation for Automatic Speaker Verification
par: Truong, Duc-Tuan, et autres
Publié: (2023)