Zero-shot Emotion Annotation in Facial Images Using Large Multimodal Models: Benchmarking and Prospects for Multi-Class, Multi-Frame Approaches
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, He, Fu, Xinyi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation
by: Zhang, He, et al.
Published: (2025)
by: Zhang, He, et al.
Published: (2025)
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
by: Si, Yonghao, et al.
Published: (2026)
by: Si, Yonghao, et al.
Published: (2026)
Learning Annotation Consensus for Continuous Emotion Recognition
by: Shoer, Ibrahim, et al.
Published: (2025)
by: Shoer, Ibrahim, et al.
Published: (2025)
Exploring Thermography Technology: A Comprehensive Facial Dataset for Face Detection, Recognition, and Emotion
by: Abuhussein, Mohamed Fawzi Abdelshafie, et al.
Published: (2024)
by: Abuhussein, Mohamed Fawzi Abdelshafie, et al.
Published: (2024)
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
by: Chen, Haodong, et al.
Published: (2024)
by: Chen, Haodong, et al.
Published: (2024)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
by: Lachmann, Tim, et al.
Published: (2026)
by: Lachmann, Tim, et al.
Published: (2026)
Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
by: Tsangko, Iosif, et al.
Published: (2025)
by: Tsangko, Iosif, et al.
Published: (2025)
Visual Affect Analysis: Predicting Emotions of Image Viewers with Vision-Language Models
by: Nowicki, Filip, et al.
Published: (2026)
by: Nowicki, Filip, et al.
Published: (2026)
A Monocular SLAM-based Multi-User Positioning System with Image Occlusion in Augmented Reality
by: Lien, Wei-Hsiang, et al.
Published: (2024)
by: Lien, Wei-Hsiang, et al.
Published: (2024)
VRMN-bD: A Multi-modal Natural Behavior Dataset of Immersive Human Fear Responses in VR Stand-up Interactive Games
by: Zhang, He, et al.
Published: (2024)
by: Zhang, He, et al.
Published: (2024)
Foundation Models for Zero-Shot Segmentation of Scientific Images without AI-Ready Data
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
by: Mukherjee, Shubhabrata, et al.
Published: (2025)
The Latency Wall: Benchmarking Off-the-Shelf Emotion Recognition for Real-Time Virtual Avatars
by: Benyamin, Yarin
Published: (2026)
by: Benyamin, Yarin
Published: (2026)
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
by: Xia, Ding, et al.
Published: (2025)
by: Xia, Ding, et al.
Published: (2025)
Has the Virtualization of the Face Changed Facial Perception? A Study of the Impact of Photo Editing and Augmented Reality on Facial Perception
by: Conwill, Louisa, et al.
Published: (2023)
by: Conwill, Louisa, et al.
Published: (2023)
Improvement in Facial Emotion Recognition using Synthetic Data Generated by Diffusion Model
by: Roy, Arnab Kumar, et al.
Published: (2024)
by: Roy, Arnab Kumar, et al.
Published: (2024)
Weak-Annotation of HAR Datasets using Vision Foundation Models
by: Bock, Marius, et al.
Published: (2024)
by: Bock, Marius, et al.
Published: (2024)
Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent Disambiguation
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
Exploring the "Great Unseen" in Medieval Manuscripts: Instance-Level Labeling of Legacy Image Collections with Zero-Shot Models
by: Meinecke, Christofer, et al.
Published: (2025)
by: Meinecke, Christofer, et al.
Published: (2025)
Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction
by: Sezer, Berk, et al.
Published: (2026)
by: Sezer, Berk, et al.
Published: (2026)
PoseDriver: A Unified Approach to Multi-Category Skeleton Detection for Autonomous Driving
by: Borhani, Yasamin, et al.
Published: (2026)
by: Borhani, Yasamin, et al.
Published: (2026)
SCHEMA for Gemini 3 Pro Image: A Structured Methodology for Controlled AI Image Generation on Google's Native Multimodal Model
by: Cazzaniga, Luca
Published: (2026)
by: Cazzaniga, Luca
Published: (2026)
When, Where, and What? A Novel Benchmark for Accident Anticipation and Localization with Large Language Models
by: Liao, Haicheng, et al.
Published: (2024)
by: Liao, Haicheng, et al.
Published: (2024)
Is Medieval Distant Viewing Possible? : Extending and Enriching Annotation of Legacy Image Collections using Visual Analytics
by: Meinecke, Christofer, et al.
Published: (2022)
by: Meinecke, Christofer, et al.
Published: (2022)
PromptArtisan: Multi-instruction Image Editing in Single Pass with Complete Attention Control
by: Swami, Kunal, et al.
Published: (2025)
by: Swami, Kunal, et al.
Published: (2025)
Human-in-the-Loop Annotation for Image-Based Engagement Estimation: Assessing the Impact of Model Reliability on Annotation Accuracy
by: Subramanya, Sahana Yadnakudige, et al.
Published: (2025)
by: Subramanya, Sahana Yadnakudige, et al.
Published: (2025)
SVFAP: Self-supervised Video Facial Affect Perceiver
by: Sun, Licai, et al.
Published: (2023)
by: Sun, Licai, et al.
Published: (2023)
Multi-Masked Querying Network for Robust Emotion Recognition from Incomplete Multi-Modal Physiological Signals
by: Xu, Geng-Xin, et al.
Published: (2025)
by: Xu, Geng-Xin, et al.
Published: (2025)
Facial Movement Dynamics Reveal Workload During Complex Multitasking
by: Sale, Carter, et al.
Published: (2026)
by: Sale, Carter, et al.
Published: (2026)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
by: Eing, Lennart, et al.
Published: (2026)
by: Eing, Lennart, et al.
Published: (2026)
Zero-shot Human Pose Estimation using Diffusion-based Inverse solvers
by: Karnoor, Sahil Bhandary, et al.
Published: (2025)
by: Karnoor, Sahil Bhandary, et al.
Published: (2025)
PrivatEyes: Appearance-based Gaze Estimation Using Federated Secure Multi-Party Computation
by: Elfares, Mayar, et al.
Published: (2024)
by: Elfares, Mayar, et al.
Published: (2024)
Viewpoint Recommendation for Point Cloud Labeling through Interaction Cost Modeling
by: Zhang, Yu, et al.
Published: (2026)
by: Zhang, Yu, et al.
Published: (2026)
Efficient Listener: Dyadic Facial Motion Synthesis via Action Diffusion
by: Wang, Zesheng, et al.
Published: (2025)
by: Wang, Zesheng, et al.
Published: (2025)
VoxelKeypointFusion: Generalizable Multi-View Multi-Person Pose Estimation
by: Bermuth, Daniel, et al.
Published: (2024)
by: Bermuth, Daniel, et al.
Published: (2024)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
PCNN: Probable-Class Nearest-Neighbor Explanations Improve Fine-Grained Image Classification Accuracy for AIs and Humans
by: Giang, et al.
Published: (2023)
by: Giang, et al.
Published: (2023)
ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
by: Cheng, Luo, et al.
Published: (2025)
by: Cheng, Luo, et al.
Published: (2025)
ImaGGen: Zero-Shot Generation of Co-Speech Semantic Gestures Grounded in Language and Image Input
by: Voss, Hendric, et al.
Published: (2025)
by: Voss, Hendric, et al.
Published: (2025)
Milmer: a Framework for Multiple Instance Learning based Multimodal Emotion Recognition
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
Similar Items
-
Augmenting Image Annotation: A Human-LMM Collaborative Framework for Efficient Object Selection and Label Generation
by: Zhang, He, et al.
Published: (2025) -
CytoCrowd: A Multi-Annotator Benchmark Dataset for Cytology Image Analysis
by: Si, Yonghao, et al.
Published: (2026) -
Learning Annotation Consensus for Continuous Emotion Recognition
by: Shoer, Ibrahim, et al.
Published: (2025) -
Exploring Thermography Technology: A Comprehensive Facial Dataset for Face Detection, Recognition, and Emotion
by: Abuhussein, Mohamed Fawzi Abdelshafie, et al.
Published: (2024) -
FineCLIPER: Multi-modal Fine-grained CLIP for Dynamic Facial Expression Recognition with AdaptERs
by: Chen, Haodong, et al.
Published: (2024)