MoTAS: MoE-Guided Feature Selection from TTS-Augmented Speech for Enhanced Multimodal Alzheimer's Early Screening
Fuente:
arXiv
Guardado en:
| Autores principales: | Shao, Yongqi, Mei, Binxin, Tan, Cong, Huo, Hong, Fang, Tao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
por: You, Fuming, et al.
Publicado: (2024)
por: You, Fuming, et al.
Publicado: (2024)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
por: Pan, Zihan, et al.
Publicado: (2025)
por: Pan, Zihan, et al.
Publicado: (2025)
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Can We Hear from Events? Generating Speech from Event Camera
por: Fang, Jingping, et al.
Publicado: (2026)
por: Fang, Jingping, et al.
Publicado: (2026)
AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
por: Wang, Yunsheng, et al.
Publicado: (2026)
por: Wang, Yunsheng, et al.
Publicado: (2026)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
por: Ning, Jinzhong, et al.
Publicado: (2025)
por: Ning, Jinzhong, et al.
Publicado: (2025)
Private Speech Classification without Collapse: Stabilized DP Training and Offline Distillation
por: Wen, Yadi, et al.
Publicado: (2026)
por: Wen, Yadi, et al.
Publicado: (2026)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
por: Peng, Yuezhang, et al.
Publicado: (2025)
por: Peng, Yuezhang, et al.
Publicado: (2025)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
por: Wang, Siyu, et al.
Publicado: (2025)
por: Wang, Siyu, et al.
Publicado: (2025)
TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation
por: Lee, Yubeen, et al.
Publicado: (2025)
por: Lee, Yubeen, et al.
Publicado: (2025)
MMED: A Multimodal Micro-Expression Dataset based on Audio-Visual Fusion
por: Wang, Junbo, et al.
Publicado: (2025)
por: Wang, Junbo, et al.
Publicado: (2025)
MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding
por: Yang, Meng, et al.
Publicado: (2026)
por: Yang, Meng, et al.
Publicado: (2026)
InconVAD: A Two-Stage Dual-Tower Framework for Multimodal Emotion Inconsistency Detection
por: Li, Zongyi, et al.
Publicado: (2025)
por: Li, Zongyi, et al.
Publicado: (2025)
MoLT: Mixture of Layer-Wise Tokens for Efficient Audio-Visual Learning
por: Rho, Kyeongha, et al.
Publicado: (2025)
por: Rho, Kyeongha, et al.
Publicado: (2025)
HDA-SELD: Hierarchical Cross-Modal Distillation with Multi-Level Data Augmentation for Low-Resource Audio-Visual Sound Event Localization and Detection
por: Wang, Qing, et al.
Publicado: (2025)
por: Wang, Qing, et al.
Publicado: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025)
por: Shi, Weiyan, et al.
Publicado: (2025)
Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition
por: Xia, Haiying, et al.
Publicado: (2025)
por: Xia, Haiying, et al.
Publicado: (2025)
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
por: Li, Jizhen, et al.
Publicado: (2024)
por: Li, Jizhen, et al.
Publicado: (2024)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
por: Ma, Ziyang, et al.
Publicado: (2026)
por: Ma, Ziyang, et al.
Publicado: (2026)
Manipulated Regions Localization For Partially Deepfake Audio: A Survey
por: He, Jiayi, et al.
Publicado: (2025)
por: He, Jiayi, et al.
Publicado: (2025)
Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
por: Xu, Xinmeng, et al.
Publicado: (2026)
por: Xu, Xinmeng, et al.
Publicado: (2026)
Anomaly Detection and Localization for Speech Deepfakes via Feature Pyramid Matching
por: Coletta, Emma, et al.
Publicado: (2025)
por: Coletta, Emma, et al.
Publicado: (2025)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
por: Su, Fei, et al.
Publicado: (2026)
por: Su, Fei, et al.
Publicado: (2026)
Research on Piano Timbre Transformation System Based on Diffusion Model
por: Hsu, Chun-Chieh, et al.
Publicado: (2026)
por: Hsu, Chun-Chieh, et al.
Publicado: (2026)
Multimodal Speech Enhancement Using Burst Propagation
por: Raza, Mohsin, et al.
Publicado: (2022)
por: Raza, Mohsin, et al.
Publicado: (2022)
Low-latency Speech Enhancement via Speech Token Generation
por: Xue, Huaying, et al.
Publicado: (2023)
por: Xue, Huaying, et al.
Publicado: (2023)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
por: Cui, Meng, et al.
Publicado: (2023)
por: Cui, Meng, et al.
Publicado: (2023)
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
por: Zhang, Yu, et al.
Publicado: (2025)
por: Zhang, Yu, et al.
Publicado: (2025)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
por: Cui, Yang, et al.
Publicado: (2025)
por: Cui, Yang, et al.
Publicado: (2025)
Constructing Composite Features for Interpretable Music-Tagging
por: Xue, Chenhao, et al.
Publicado: (2026)
por: Xue, Chenhao, et al.
Publicado: (2026)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
por: Lu, Wenhuan, et al.
Publicado: (2025)
por: Lu, Wenhuan, et al.
Publicado: (2025)
A Survey on Cross-Modal Interaction Between Music and Multimodal Data
por: Li, Sifei, et al.
Publicado: (2025)
por: Li, Sifei, et al.
Publicado: (2025)
MLLM-based Speech Recognition: When and How is Multimodality Beneficial?
por: Guan, Yiwen, et al.
Publicado: (2025)
por: Guan, Yiwen, et al.
Publicado: (2025)
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction
por: Zhao, Yuan, et al.
Publicado: (2024)
por: Zhao, Yuan, et al.
Publicado: (2024)
Preserving Speaker Information in Direct Speech-to-Speech Translation with Non-Autoregressive Generation and Pretraining
por: Zhou, Rui, et al.
Publicado: (2024)
por: Zhou, Rui, et al.
Publicado: (2024)
AUREXA-SE: Audio-Visual Unified Representation Exchange Architecture with Cross-Attention and Squeezeformer for Speech Enhancement
por: Sajid, M., et al.
Publicado: (2025)
por: Sajid, M., et al.
Publicado: (2025)
Conformer-based Ultrasound-to-Speech Conversion
por: Ibrahimov, Ibrahim, et al.
Publicado: (2025)
por: Ibrahimov, Ibrahim, et al.
Publicado: (2025)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
por: Herbuela, Von Ralph Dane Marquez, et al.
Publicado: (2025)
Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition
por: Koo, Inyong, et al.
Publicado: (2026)
por: Koo, Inyong, et al.
Publicado: (2026)
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
por: Bang, Hayeon, et al.
Publicado: (2025)
por: Bang, Hayeon, et al.
Publicado: (2025)
Ejemplares similares
-
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
por: You, Fuming, et al.
Publicado: (2024) -
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
por: Pan, Zihan, et al.
Publicado: (2025) -
MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions
por: Li, Junjie, et al.
Publicado: (2025) -
Can We Hear from Events? Generating Speech from Event Camera
por: Fang, Jingping, et al.
Publicado: (2026) -
AMB-DSGDN: Adaptive Modality-Balanced Dynamic Semantic Graph Differential Network for Multimodal Emotion Recognition
por: Wang, Yunsheng, et al.
Publicado: (2026)