Nuanced Emotion Recognition Based on a Segment-based MLLM Framework Leveraging Qwen3-Omni for AH Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Tang, Liang, Li, Hongda, Zhang, Jiayu, Chen, Long, Li, Shuxian, Pei, Siqi, Duan, Tiaonan, Cheng, Yuhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
by: Cheng, Yuhao, et al.
Published: (2025)
by: Cheng, Yuhao, et al.
Published: (2025)
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026)
by: Pei, Siqi, et al.
Published: (2026)
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025)
by: Huang, Dawei, et al.
Published: (2025)
Qwen3-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
by: Xuan, Shiyu, et al.
Published: (2026)
by: Xuan, Shiyu, et al.
Published: (2026)
Qwen3.5-Omni Technical Report
by: Qwen Team
Published: (2026)
by: Qwen Team
Published: (2026)
Qwen2.5-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025)
by: Xu, Jin, et al.
Published: (2025)
BeMERC: Behavior-Aware MLLM-based Framework for Multimodal Emotion Recognition in Conversation
by: Fu, Yumeng, et al.
Published: (2025)
by: Fu, Yumeng, et al.
Published: (2025)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
by: Chao, Yuhao, et al.
Published: (2025)
by: Chao, Yuhao, et al.
Published: (2025)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
by: Wu, Zehui, et al.
Published: (2024)
by: Wu, Zehui, et al.
Published: (2024)
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
by: Chang, Jingjing, et al.
Published: (2025)
by: Chang, Jingjing, et al.
Published: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
by: Li, Mingxin, et al.
Published: (2026)
by: Li, Mingxin, et al.
Published: (2026)
CAND: Cross-Domain Ambiguity Inference for Early Detecting Nuanced Illness Deterioration
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
by: Ting, Lo Pang-Yun, et al.
Published: (2025)
Phishing Detection in Ethereum via Temporal Graph Contrastive Learning
by: Wu, Cong, et al.
Published: (2026)
by: Wu, Cong, et al.
Published: (2026)
OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
EERPD: Leveraging Emotion and Emotion Regulation for Improving Personality Detection
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition
by: Liu, Yu, et al.
Published: (2025)
by: Liu, Yu, et al.
Published: (2025)
Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
by: Wu, Diankun, et al.
Published: (2025)
by: Wu, Diankun, et al.
Published: (2025)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
by: Wang, Hsuan-Yu, et al.
Published: (2025)
by: Wang, Hsuan-Yu, et al.
Published: (2025)
SkeletonAgent: An Agentic Interaction Framework for Skeleton-based Action Recognition
by: Liu, Hongda, et al.
Published: (2025)
by: Liu, Hongda, et al.
Published: (2025)
Omni-Fusion of Spatial and Spectral for Hyperspectral Image Segmentation
by: Zhang, Qing, et al.
Published: (2025)
by: Zhang, Qing, et al.
Published: (2025)
Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
by: Fan, Qi, et al.
Published: (2024)
by: Fan, Qi, et al.
Published: (2024)
Qwen3 Technical Report
by: Yang, An, et al.
Published: (2025)
by: Yang, An, et al.
Published: (2025)
Qwen-Image Technical Report
by: Wu, Chenfei, et al.
Published: (2025)
by: Wu, Chenfei, et al.
Published: (2025)
MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
by: Zhang, Yiyi, et al.
Published: (2025)
by: Zhang, Yiyi, et al.
Published: (2025)
Federated Dialogue-Semantic Diffusion for Emotion Recognition under Incomplete Modalities
by: Qiu, Xihang, et al.
Published: (2025)
by: Qiu, Xihang, et al.
Published: (2025)
Qwen3Guard Technical Report
by: Zhao, Haiquan, et al.
Published: (2025)
by: Zhao, Haiquan, et al.
Published: (2025)
Relaxed Weak Accelerated Proximal Gradient Method: a Unified Framework for Nesterov's Accelerations
by: Li, Hongda, et al.
Published: (2025)
by: Li, Hongda, et al.
Published: (2025)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
by: Zou, Yueying, et al.
Published: (2025)
by: Zou, Yueying, et al.
Published: (2025)
EmoMM: Benchmarking and Steering MLLM for Multimodal Emotion Recognition under Conflict and Missingness
by: Sun, Yueru, et al.
Published: (2026)
by: Sun, Yueru, et al.
Published: (2026)
Moment and Highlight Detection via MLLM Frame Segmentation
by: Jiwanta, I Putu Andika Bagas, et al.
Published: (2025)
by: Jiwanta, I Putu Andika Bagas, et al.
Published: (2025)
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Qwen2 Technical Report
by: Yang, An, et al.
Published: (2024)
by: Yang, An, et al.
Published: (2024)
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
by: Lyu, Xiaosen, et al.
Published: (2025)
by: Lyu, Xiaosen, et al.
Published: (2025)
Similar Items
-
Evolving in Tasks: Empowering the Multi-modality Large Language Model as the Computer Use Agent
by: Cheng, Yuhao, et al.
Published: (2025) -
AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction Refinement
by: Pei, Siqi, et al.
Published: (2026) -
Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
by: Huang, Dawei, et al.
Published: (2025) -
Qwen3-Omni Technical Report
by: Xu, Jin, et al.
Published: (2025) -
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
by: Xuan, Shiyu, et al.
Published: (2026)