Facial Expression Recognition with YOLOv11 and YOLOv12: A Comparative Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Aymon, Umma, Kamarudin, Nur Shazwani, Nasir, Ahmad Fakhri Ab.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911265745534976
author Aymon, Umma
Kamarudin, Nur Shazwani
Nasir, Ahmad Fakhri Ab.
author_facet Aymon, Umma
Kamarudin, Nur Shazwani
Nasir, Ahmad Fakhri Ab.
contents Facial Expression Recognition remains a challenging task, especially in unconstrained, real-world environments. This study investigates the performance of two lightweight models, YOLOv11n and YOLOv12n, which are the nano variants of the latest official YOLO series, within a unified detection and classification framework for FER. Two benchmark classification datasets, FER2013 and KDEF, are converted into object detection format and model performance is evaluated using mAP 0.5, precision, recall, and confusion matrices. Results show that YOLOv12n achieves the highest overall performance on the clean KDEF dataset with a mAP 0.5 of 95.6, and also outperforms YOLOv11n on the FER2013 dataset in terms of mAP 63.8, reflecting stronger sensitivity to varied expressions. In contrast, YOLOv11n demonstrates higher precision 65.2 on FER2013, indicating fewer false positives and better reliability in noisy, real-world conditions. On FER2013, both models show more confusion between visually similar expressions, while clearer class separation is observed on the cleaner KDEF dataset. These findings underscore the trade-off between sensitivity and precision, illustrating how lightweight YOLO models can effectively balance performance and efficiency. The results demonstrate adaptability across both controlled and real-world conditions, establishing these models as strong candidates for real-time, resource-constrained emotion-aware AI applications.
format Preprint
id arxiv_https___arxiv_org_abs_2511_10940
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Facial Expression Recognition with YOLOv11 and YOLOv12: A Comparative Study
Aymon, Umma
Kamarudin, Nur Shazwani
Nasir, Ahmad Fakhri Ab.
Computer Vision and Pattern Recognition
Facial Expression Recognition remains a challenging task, especially in unconstrained, real-world environments. This study investigates the performance of two lightweight models, YOLOv11n and YOLOv12n, which are the nano variants of the latest official YOLO series, within a unified detection and classification framework for FER. Two benchmark classification datasets, FER2013 and KDEF, are converted into object detection format and model performance is evaluated using mAP 0.5, precision, recall, and confusion matrices. Results show that YOLOv12n achieves the highest overall performance on the clean KDEF dataset with a mAP 0.5 of 95.6, and also outperforms YOLOv11n on the FER2013 dataset in terms of mAP 63.8, reflecting stronger sensitivity to varied expressions. In contrast, YOLOv11n demonstrates higher precision 65.2 on FER2013, indicating fewer false positives and better reliability in noisy, real-world conditions. On FER2013, both models show more confusion between visually similar expressions, while clearer class separation is observed on the cleaner KDEF dataset. These findings underscore the trade-off between sensitivity and precision, illustrating how lightweight YOLO models can effectively balance performance and efficiency. The results demonstrate adaptability across both controlled and real-world conditions, establishing these models as strong candidates for real-time, resource-constrained emotion-aware AI applications.
title Facial Expression Recognition with YOLOv11 and YOLOv12: A Comparative Study
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.10940