Topological Deep Learning for Speech Data
Fuente:
arXiv
Saved in:
| Main Author: | Yu, Zhiwang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
by: Neururer, Daniel, et al.
Published: (2023)
by: Neururer, Daniel, et al.
Published: (2023)
Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization
by: Raghav, Nikhil, et al.
Published: (2024)
by: Raghav, Nikhil, et al.
Published: (2024)
DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
by: Mu, Zhaoxi, et al.
Published: (2024)
by: Mu, Zhaoxi, et al.
Published: (2024)
FairSSD: Understanding Bias in Synthetic Speech Detectors
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
Benchmarking Machine Learning Methods for Distributed Acoustic Sensing
by: Shi, Shuaikai, et al.
Published: (2025)
by: Shi, Shuaikai, et al.
Published: (2025)
Characterizing Continual Learning Scenarios and Strategies for Audio Analysis
by: Bhatt, Ruchi, et al.
Published: (2024)
by: Bhatt, Ruchi, et al.
Published: (2024)
Automated Detection of Dolphin Whistles with Convolutional Networks and Transfer Learning
by: Korkmaz, Burla Nur, et al.
Published: (2022)
by: Korkmaz, Burla Nur, et al.
Published: (2022)
Compression Robust Synthetic Speech Detection Using Patched Spectrogram Transformer
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)
Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
by: Rehman, Yasar Abbas Ur, et al.
Published: (2024)
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
by: Airale, Louis, et al.
Published: (2023)
by: Airale, Louis, et al.
Published: (2023)
A Study of Dropout-Induced Modality Bias on Robustness to Missing Video Frames for Audio-Visual Speech Recognition
by: Dai, Yusheng, et al.
Published: (2024)
by: Dai, Yusheng, et al.
Published: (2024)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
Spiking Structured State Space Model for Monaural Speech Enhancement
by: Du, Yu, et al.
Published: (2023)
by: Du, Yu, et al.
Published: (2023)
JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
by: Zhang, Haomin, et al.
Published: (2025)
by: Zhang, Haomin, et al.
Published: (2025)
MMAudioSep: Taming Video-to-Audio Generative Model Towards Video/Text-Queried Sound Separation
by: Takahashi, Akira, et al.
Published: (2025)
by: Takahashi, Akira, et al.
Published: (2025)
Training-Free Voice Conversion with Factorized Optimal Transport
by: Lobashev, Alexander, et al.
Published: (2025)
by: Lobashev, Alexander, et al.
Published: (2025)
Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
by: Hayakawa, Akio, et al.
Published: (2025)
by: Hayakawa, Akio, et al.
Published: (2025)
Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
by: Zeng, Runhao, et al.
Published: (2025)
by: Zeng, Runhao, et al.
Published: (2025)
Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
by: Li, Tianqi, et al.
Published: (2024)
by: Li, Tianqi, et al.
Published: (2024)
Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
Acoustic Scene Classification: A Competition Review
by: Gharib, Shayan, et al.
Published: (2018)
by: Gharib, Shayan, et al.
Published: (2018)
Exploring Green AI for Audio Deepfake Detection
by: Saha, Subhajit, et al.
Published: (2024)
by: Saha, Subhajit, et al.
Published: (2024)
Joint Multimodal Transformer for Emotion Recognition in the Wild
by: Waligora, Paul, et al.
Published: (2024)
by: Waligora, Paul, et al.
Published: (2024)
Dynamic Cross Attention for Audio-Visual Person Verification
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Dynamic Modality and View Selection for Multimodal Emotion Recognition with Missing Modalities
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
by: Menon, Luciana Trinkaus, et al.
Published: (2024)
An Eye for an Ear: Zero-shot Audio Description Leveraging an Image Captioner using Audiovisual Distribution Alignment
by: Malard, Hugo, et al.
Published: (2024)
by: Malard, Hugo, et al.
Published: (2024)
Developing an AI-based Integrated System for Bee Health Evaluation
by: Liang, Andrew
Published: (2024)
by: Liang, Andrew
Published: (2024)
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
by: Cheng, Luyao, et al.
Published: (2024)
by: Cheng, Luyao, et al.
Published: (2024)
The Solution for Temporal Sound Localisation Task of ICCV 1st Perception Test Challenge 2023
by: Huang, Yurui, et al.
Published: (2024)
by: Huang, Yurui, et al.
Published: (2024)
Character-aware audio-visual subtitling in context
by: Huh, Jaesung, et al.
Published: (2024)
by: Huh, Jaesung, et al.
Published: (2024)
Towards reliable respiratory disease diagnosis based on cough sounds and vision transformers
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
A High-Accuracy Optical Music Recognition Method Based on Bottleneck Residual Convolutions
by: Ma, Junwen, et al.
Published: (2026)
by: Ma, Junwen, et al.
Published: (2026)
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
by: Cheng, Ho Kei, et al.
Published: (2024)
by: Cheng, Ho Kei, et al.
Published: (2024)
Attention Isn't All You Need for Emotion Recognition:Domain Features Outperform Transformers on the EAV Dataset
by: Guragain, Anmol
Published: (2026)
by: Guragain, Anmol
Published: (2026)
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
by: Haji-Ali, Moayed, et al.
Published: (2024)
by: Haji-Ali, Moayed, et al.
Published: (2024)
SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound
by: Dagli, Rishit, et al.
Published: (2024)
by: Dagli, Rishit, et al.
Published: (2024)
Similar Items
-
Deep Neural Networks for Automatic Speaker Recognition Do Not Learn Supra-Segmental Temporal Features
by: Neururer, Daniel, et al.
Published: (2023) -
Assessing the Robustness of Spectral Clustering for Deep Speaker Diarization
by: Raghav, Nikhil, et al.
Published: (2024) -
DiVISe: Direct Visual-Input Speech Synthesis Preserving Speaker Characteristics And Intelligibility
by: Liu, Yifan, et al.
Published: (2025) -
Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction
by: Mu, Zhaoxi, et al.
Published: (2024) -
FairSSD: Understanding Bias in Synthetic Speech Detectors
by: Yadav, Amit Kumar Singh, et al.
Published: (2024)