Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Lipeng, Xiong, Yifan, Guo, Dongyue, Mo, Wei, Yu, Lingyu, Yang, Hui, Lin, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024)
von: Lin, Jinwei
Veröffentlicht: (2024)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
von: Lin, Yueqian, et al.
Veröffentlicht: (2025)
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
SVDD 2024: The Inaugural Singing Voice Deepfake Detection Challenge
von: Zhang, You, et al.
Veröffentlicht: (2024)
von: Zhang, You, et al.
Veröffentlicht: (2024)
MuseAgent-1: Interactive Grounded Multimodal Understanding of Music Scores and Performance Audio
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
von: Zhao, Qihao, et al.
Veröffentlicht: (2026)
HybridVC: Efficient Voice Style Conversion with Text and Audio Prompts
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
von: Niu, Xinlei, et al.
Veröffentlicht: (2024)
Towards Source Attribution of Singing Voice Deepfake with Multimodal Foundation Models
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
von: Clarke, Jason, et al.
Veröffentlicht: (2025)
RVCBench: Benchmarking the Robustness of Voice Cloning Across Modern Audio Generation Models
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
von: Jin, Ruinan, et al.
Veröffentlicht: (2026)
PerformSinger: Multimodal Singing Voice Synthesis Leveraging Synchronized Lip Cues from Singing Performance Videos
von: Gu, Ke, et al.
Veröffentlicht: (2025)
von: Gu, Ke, et al.
Veröffentlicht: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Optimizing Feature Extraction for Symbolic Music
von: Simonetta, Federico, et al.
Veröffentlicht: (2023)
von: Simonetta, Federico, et al.
Veröffentlicht: (2023)
Dance2MIDI: Dance-driven multi-instruments music generation
von: Han, Bo, et al.
Veröffentlicht: (2023)
von: Han, Bo, et al.
Veröffentlicht: (2023)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
von: Jiang, Ziyang, et al.
Veröffentlicht: (2024)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
von: Huang, Lian, et al.
Veröffentlicht: (2024)
von: Huang, Lian, et al.
Veröffentlicht: (2024)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
von: Li, Jizhen, et al.
Veröffentlicht: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
A Study on Synthesizing Expressive Violin Performances: Approaches and Comparisons
von: Hung, Tzu-Yun, et al.
Veröffentlicht: (2024)
von: Hung, Tzu-Yun, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
FastTalker: Jointly Generating Speech and Conversational Gestures from Text
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
von: Guo, Zixin, et al.
Veröffentlicht: (2024)
CommonVoice-SpeechRE and RPG-MoGe: Advancing Speech Relation Extraction with a New Dataset and Multi-Order Generative Framework
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
von: Ning, Jinzhong, et al.
Veröffentlicht: (2025)
Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
Resource-Efficient Reference-Free Evaluation of Audio Captions
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
von: Mahfuz, Rehana, et al.
Veröffentlicht: (2024)
Song Aesthetics Evaluation with Multi-Stem Attention and Hierarchical Uncertainty Modeling
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
von: Lv, Yishan, et al.
Veröffentlicht: (2026)
MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
von: Wang, Yutian, et al.
Veröffentlicht: (2024)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
Multi-Track MusicLDM: Towards Versatile Music Generation with Latent Diffusion Model
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
M$^{2}$UGen: Multi-modal Music Understanding and Generation with the Power of Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
von: Liu, Shansong, et al.
Veröffentlicht: (2023)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
von: Su, Yi, et al.
Veröffentlicht: (2025)
von: Su, Yi, et al.
Veröffentlicht: (2025)
Visual-based spatial audio generation system for multi-speaker environments
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2025)
Improved symbolic drum style classification with grammar-based hierarchical representations
von: Géré, Léo, et al.
Veröffentlicht: (2024)
von: Géré, Léo, et al.
Veröffentlicht: (2024)
MuMu-LLaMA: Multi-modal Music Understanding and Generation via Large Language Models
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
von: Liu, Shansong, et al.
Veröffentlicht: (2024)
pyAMPACT: A Score-Audio Alignment Toolkit for Performance Data Estimation and Multi-modal Processing
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
von: Devaney, Johanna, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
von: Lin, Jinwei
Veröffentlicht: (2024) -
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025) -
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025) -
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
von: Lin, Yueqian, et al.
Veröffentlicht: (2025) -
CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)