Beyond Classification: Towards Speech Emotion Reasoning with Multitask AudioLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wenyu, He, Yingxu, Lin, Geyu, Liu, Zhuohan, Sun, Shuo, Wang, Bin, Zou, Xunlong, Wong, Jeremy H. M., Wang, Qiongqiong, Sailor, Hardik B., Chen, Nancy F., Aw, Ai Ti |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
by: Zhang, Wenyu, et al.
Published: (2024)
by: Zhang, Wenyu, et al.
Published: (2024)
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
by: He, Yingxu, et al.
Published: (2024)
by: He, Yingxu, et al.
Published: (2024)
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
by: Wang, Qiongqiong, et al.
Published: (2025)
by: Wang, Qiongqiong, et al.
Published: (2025)
AudioBench: A Universal Benchmark for Audio Large Language Models
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
by: Wang, Qiongqiong, et al.
Published: (2025)
by: Wang, Qiongqiong, et al.
Published: (2025)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
by: Wang, Qiongqiong, et al.
Published: (2025)
by: Wang, Qiongqiong, et al.
Published: (2025)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
by: Huzaifah, Muhammad, et al.
Published: (2024)
by: Huzaifah, Muhammad, et al.
Published: (2024)
Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models
by: Wang, Bin, et al.
Published: (2025)
by: Wang, Bin, et al.
Published: (2025)
Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs
by: Quang, Trung Nguyen, et al.
Published: (2026)
by: Quang, Trung Nguyen, et al.
Published: (2026)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
by: Pan, Zihan, et al.
Published: (2024)
by: Pan, Zihan, et al.
Published: (2024)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
by: Guragain, Anmol, et al.
Published: (2024)
by: Guragain, Anmol, et al.
Published: (2024)
Beyond Transcription: Unified Audio Schema for Perception-Aware AudioLLMs
by: Zhang, Linhao, et al.
Published: (2026)
by: Zhang, Linhao, et al.
Published: (2026)
IFEval-Audio: Benchmarking Instruction-Following Capability in Audio-based Large Language Models
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
by: Sailor, Hardik B., et al.
Published: (2025)
by: Sailor, Hardik B., et al.
Published: (2025)
MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection
by: Pan, Zihan, et al.
Published: (2025)
by: Pan, Zihan, et al.
Published: (2025)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
by: Liu, Tianchi, et al.
Published: (2024)
by: Liu, Tianchi, et al.
Published: (2024)
Unlocking Cognitive Capabilities and Analyzing the Perception-Logic Trade-off
by: Zhang, Longyin, et al.
Published: (2026)
by: Zhang, Longyin, et al.
Published: (2026)
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Improving Multilingual Social Media Insights: Aspect-based Comment Analysis
by: Zhang, Longyin, et al.
Published: (2025)
by: Zhang, Longyin, et al.
Published: (2025)
AR&D: A Framework for Retrieving and Describing Concepts for Interpreting AudioLLMs
by: Chowdhury, Townim Faisal, et al.
Published: (2026)
by: Chowdhury, Townim Faisal, et al.
Published: (2026)
Semi-supervised Learning For Robust Speech Evaluation
by: Zhang, Huayun, et al.
Published: (2024)
by: Zhang, Huayun, et al.
Published: (2024)
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
by: Shahin, Mostafa, et al.
Published: (2025)
by: Shahin, Mostafa, et al.
Published: (2025)
Interpolating Speaker Identities in Embedding Space for Data Expansion
by: Liu, Tianchi, et al.
Published: (2025)
by: Liu, Tianchi, et al.
Published: (2025)
Quantizer-Aware Hierarchical Neural Codec Modeling for Speech Deepfake Detection
by: Wu, Jinyang, et al.
Published: (2026)
by: Wu, Jinyang, et al.
Published: (2026)
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic AudioLLMs
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
by: Bhatti, Hunzalah Hassan, et al.
Published: (2026)
SEA-Spoof: Bridging The Gap in Multilingual Audio Deepfake Detection for South-East Asian
by: Wu, Jinyang, et al.
Published: (2025)
by: Wu, Jinyang, et al.
Published: (2025)
MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs
by: Ali, Zien Sheikh, et al.
Published: (2026)
by: Ali, Zien Sheikh, et al.
Published: (2026)
SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning
by: Wang, Bin, et al.
Published: (2023)
by: Wang, Bin, et al.
Published: (2023)
Opinion dynamics on biased dynamical networks: beyond rare opinion updating
by: Wang, Xunlong, et al.
Published: (2024)
by: Wang, Xunlong, et al.
Published: (2024)
CCL-XCoT: An Efficient Cross-Lingual Knowledge Transfer Method for Mitigating Hallucination Generation
by: Zheng, Weihua, et al.
Published: (2025)
by: Zheng, Weihua, et al.
Published: (2025)
Adversarial synthesis based data-augmentation for code-switched spoken language identification
by: Shastri, Parth, et al.
Published: (2022)
by: Shastri, Parth, et al.
Published: (2022)
Nonlinear contagion dynamics on dynamical networks: exact solutions ranging from consensus times to evolutionary trajectories
by: Wang, Xunlong, et al.
Published: (2025)
by: Wang, Xunlong, et al.
Published: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
by: Ritter-Gutierrez, Fabian, et al.
Published: (2024)
by: Ritter-Gutierrez, Fabian, et al.
Published: (2024)
California Research Institute on the Integration of Students with Severe Disabilities. Final Report, Years 1987-1992.
by: Sailor, Wayne, et al.
Published: (1993)
by: Sailor, Wayne, et al.
Published: (1993)
CrossIn: An Efficient Instruction Tuning Approach for Cross-Lingual Knowledge Alignment
by: Lin, Geyu, et al.
Published: (2024)
by: Lin, Geyu, et al.
Published: (2024)
M4SER: Multimodal, Multirepresentation, Multitask, and Multistrategy Learning for Speech Emotion Recognition
by: He, Jiajun, et al.
Published: (2025)
by: He, Jiajun, et al.
Published: (2025)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
by: Wu, Zehui, et al.
Published: (2024)
by: Wu, Zehui, et al.
Published: (2024)
Telmisartan Ameliorates Hypertension‐Induced Kidney Damage by Restoring Glomerular Endothelial Barrier Function
by: Qiongqiong Guo, et al.
Published: (2025)
by: Qiongqiong Guo, et al.
Published: (2025)
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
by: Wang, Sirui, et al.
Published: (2025)
by: Wang, Sirui, et al.
Published: (2025)
DFALLM: Achieving Generalizable Multitask Deepfake Detection by Optimizing Audio LLM Components
by: Li, Yupei, et al.
Published: (2025)
by: Li, Yupei, et al.
Published: (2025)
Similar Items
-
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
by: Zhang, Wenyu, et al.
Published: (2024) -
MERaLiON-AudioLLM: Bridging Audio and Language with Large Language Models
by: He, Yingxu, et al.
Published: (2024) -
Contextual Paralinguistic Data Creation for Multi-Modal Speech-LLM: Data Condensation and Spoken QA Generation
by: Wang, Qiongqiong, et al.
Published: (2025) -
AudioBench: A Universal Benchmark for Audio Large Language Models
by: Wang, Bin, et al.
Published: (2024) -
Incorporating Contextual Paralinguistic Understanding in Large Speech-Language Models
by: Wang, Qiongqiong, et al.
Published: (2025)