Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shahin, Mostafa, Ahmed, Beena, Epps, Julien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
von: Chen, Yiming, et al.
Veröffentlicht: (2024)
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
Audio-FLAN: A Preliminary Release
von: Xue, Liumeng, et al.
Veröffentlicht: (2025)
von: Xue, Liumeng, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations
von: Xu, David
Veröffentlicht: (2024)
von: Xu, David
Veröffentlicht: (2024)
Listening and Seeing Again: Generative Error Correction for Audio-Visual Speech Recognition
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
von: Rho, Kyeongha, et al.
Veröffentlicht: (2025)
DeepFake Doctor: Diagnosing and Treating Audio-Video Fake Detection
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
von: Klemt, Marcel, et al.
Veröffentlicht: (2025)
Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
von: He, Jianfeng, et al.
Veröffentlicht: (2023)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
von: Xie, Zhifei, et al.
Veröffentlicht: (2025)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
Double Mixture: Towards Continual Event Detection from Speech
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
von: Kang, Jingqi, et al.
Veröffentlicht: (2024)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
Kimi-Audio Technical Report
von: KimiTeam, et al.
Veröffentlicht: (2025)
von: KimiTeam, et al.
Veröffentlicht: (2025)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
von: Li, Maomao, et al.
Veröffentlicht: (2026)
von: Li, Maomao, et al.
Veröffentlicht: (2026)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
MM-Sonate: Multimodal Controllable Audio-Video Generation with Zero-Shot Voice Cloning
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2026)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
von: Ma, Ziyang, et al.
Veröffentlicht: (2025)
Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Neural Style Transfer for Audio Spectograms
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
von: Verma, Prateek, et al.
Veröffentlicht: (2018)
Ähnliche Einträge
-
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
von: Li, Yuanchao, et al.
Veröffentlicht: (2024) -
Beyond Single-Audio: Advancing Multi-Audio Processing in Audio Large Language Models
von: Chen, Yiming, et al.
Veröffentlicht: (2024) -
ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025) -
Audio-FLAN: A Preliminary Release
von: Xue, Liumeng, et al.
Veröffentlicht: (2025) -
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)