Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Licai, Jiang, Xingxun, Chen, Haoyu, Li, Yante, Lian, Zheng, Liu, Biu, Zong, Yuan, Zheng, Wenming, Leppänen, Jukka M., Zhao, Guoying |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Is Micro-expression Ethnic Leaning?
by: Khor, Huai-Qian, et al.
Published: (2025)
by: Khor, Huai-Qian, et al.
Published: (2025)
Infused Suppression Of Magnification Artefacts For Micro-AU Detection
by: Khor, Huai-Qian, et al.
Published: (2025)
by: Khor, Huai-Qian, et al.
Published: (2025)
Towards Consistent and Controllable Image Synthesis for Face Editing
by: Wei, Mengting, et al.
Published: (2025)
by: Wei, Mengting, et al.
Published: (2025)
MagicFace: High-Fidelity Facial Expression Editing with Action-Unit Control
by: Wei, Mengting, et al.
Published: (2025)
by: Wei, Mengting, et al.
Published: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
by: Qi, Tianhua, et al.
Published: (2024)
by: Qi, Tianhua, et al.
Published: (2024)
Deep Change Monitoring: A Hyperbolic Representative Learning Framework and a Dataset for Long-term Fine-grained Tree Change Detection
by: Li, Yante, et al.
Published: (2025)
by: Li, Yante, et al.
Published: (2025)
HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
by: Sun, Licai, et al.
Published: (2024)
by: Sun, Licai, et al.
Published: (2024)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
by: Lu, Cheng, et al.
Published: (2024)
by: Lu, Cheng, et al.
Published: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
by: Qi, Tianhua, et al.
Published: (2026)
by: Qi, Tianhua, et al.
Published: (2026)
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
by: Qi, Tianhua, et al.
Published: (2024)
by: Qi, Tianhua, et al.
Published: (2024)
Towards Localized Fine-Grained Control for Facial Expression Generation
by: Varanka, Tuomas, et al.
Published: (2024)
by: Varanka, Tuomas, et al.
Published: (2024)
Micro-AU CLIP: Fine-Grained Contrastive Learning from Local Independence to Global Dependency for Micro-Expression Action Unit Detection
by: Wei, Jinsheng, et al.
Published: (2026)
by: Wei, Jinsheng, et al.
Published: (2026)
Speech Swin-Transformer: Exploring a Hierarchical Transformer with Shifted Windows for Speech Emotion Recognition
by: Wang, Yong, et al.
Published: (2024)
by: Wang, Yong, et al.
Published: (2024)
AffectGPT: Dataset and Framework for Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric Guidance
by: Wei, Mengting, et al.
Published: (2025)
by: Wei, Mengting, et al.
Published: (2025)
Temporal Label Hierachical Network for Compound Emotion Recognition
by: Li, Sunan, et al.
Published: (2024)
by: Li, Sunan, et al.
Published: (2024)
SVFAP: Self-supervised Video Facial Affect Perceiver
by: Sun, Licai, et al.
Published: (2023)
by: Sun, Licai, et al.
Published: (2023)
EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
by: Lian, Zheng, et al.
Published: (2025)
by: Lian, Zheng, et al.
Published: (2025)
GPT-4V with Emotion: A Zero-shot Benchmark for Generalized Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
by: Zhao, Yan, et al.
Published: (2024)
by: Zhao, Yan, et al.
Published: (2024)
AffectGPT: A New Dataset, Model, and Benchmark for Emotion Understanding with Multimodal Large Language Models
by: Lian, Zheng, et al.
Published: (2025)
by: Lian, Zheng, et al.
Published: (2025)
ITEACH-Net: Inverted Teacher-studEnt seArCH Network for Emotion Recognition in Conversation
by: Sun, Haiyang, et al.
Published: (2023)
by: Sun, Haiyang, et al.
Published: (2023)
MERBench: A Unified Evaluation Benchmark for Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
FEALLM: Advancing Facial Emotion Analysis in Multimodal Large Language Models with Emotional Synergy and Reasoning
by: Hu, Zhuozhao, et al.
Published: (2025)
by: Hu, Zhuozhao, et al.
Published: (2025)
Decoupled Doubly Contrastive Learning for Cross Domain Facial Action Unit Detection
by: Li, Yong, et al.
Published: (2025)
by: Li, Yong, et al.
Published: (2025)
Axisymmetric Coil Winding Surfaces for Non-Axisymmetric Fusion Devices
by: Biu, J., et al.
Published: (2025)
by: Biu, J., et al.
Published: (2025)
Biometric Authentication Based on Enhanced Remote Photoplethysmography Signal Morphology
by: Sun, Zhaodong, et al.
Published: (2024)
by: Sun, Zhaodong, et al.
Published: (2024)
Towards Robust 3D Pose Transfer with Adversarial Learning
by: Chen, Haoyu, et al.
Published: (2024)
by: Chen, Haoyu, et al.
Published: (2024)
Multimodal Functional Maximum Correlation for Emotion Recognition
by: Zheng, Deyang, et al.
Published: (2025)
by: Zheng, Deyang, et al.
Published: (2025)
Hybrid-supervised Hypergraph-enhanced Transformer for Micro-gesture Based Emotion Recognition
by: Xia, Zhaoqiang, et al.
Published: (2025)
by: Xia, Zhaoqiang, et al.
Published: (2025)
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses
by: Xu, Shuhao, et al.
Published: (2026)
by: Xu, Shuhao, et al.
Published: (2026)
PC-MNet: Dual-Level Congruity Modeling for Multimodal Sarcasm Detection via Polarity-Modulated Attention
by: Li, Maoheng, et al.
Published: (2026)
by: Li, Maoheng, et al.
Published: (2026)
Incorporating Scene Context and Semantic Labels for Enhanced Group-level Emotion Recognition
by: Zhu, Qing, et al.
Published: (2025)
by: Zhu, Qing, et al.
Published: (2025)
Self-Supervised Facial Representation Learning with Facial Region Awareness
by: Gao, Zheng, et al.
Published: (2024)
by: Gao, Zheng, et al.
Published: (2024)
MER 2024: Semi-Supervised Learning, Noise Robustness, and Open-Vocabulary Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
SPOLRE: Semantic Preserving Object Layout Reconstruction for Image Captioning System Testing
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
How Large Language Models Are Changing MOOC Essay Answers: A Comparison of Pre- and Post-LLM Responses
by: Leppänen, Leo, et al.
Published: (2025)
by: Leppänen, Leo, et al.
Published: (2025)
Capturing Rich Behavior Representations: A Dynamic Action Semantic-Aware Graph Transformer for Video Captioning
by: Liu, Caihua, et al.
Published: (2025)
by: Liu, Caihua, et al.
Published: (2025)
OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2024)
by: Lian, Zheng, et al.
Published: (2024)
Enhanced Algorithmic Perfect State Transfer on IBM Quantum Computers
by: Ge, Zong-Yuan, et al.
Published: (2025)
by: Ge, Zong-Yuan, et al.
Published: (2025)
Similar Items
-
Is Micro-expression Ethnic Leaning?
by: Khor, Huai-Qian, et al.
Published: (2025) -
Infused Suppression Of Magnification Artefacts For Micro-AU Detection
by: Khor, Huai-Qian, et al.
Published: (2025) -
Towards Consistent and Controllable Image Synthesis for Face Editing
by: Wei, Mengting, et al.
Published: (2025) -
MagicFace: High-Fidelity Facial Expression Editing with Action-Unit Control
by: Wei, Mengting, et al.
Published: (2025) -
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
by: Qi, Tianhua, et al.
Published: (2024)