Evaluating and Predicting Distorted Human Body Parts for Generated Images

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ma, Lu, Cao, Kaibo, Liang, Hao, Lin, Jiaxin, Li, Zhuang, Liu, Yuhong, Zhang, Jihong, Zhang, Wentao, Cui, Bin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916638935220224
author Ma, Lu
Cao, Kaibo
Liang, Hao
Lin, Jiaxin
Li, Zhuang
Liu, Yuhong
Zhang, Jihong
Zhang, Wentao
Cui, Bin
author_facet Ma, Lu
Cao, Kaibo
Liang, Hao
Lin, Jiaxin
Li, Zhuang
Liu, Yuhong
Zhang, Jihong
Zhang, Wentao
Cui, Bin
contents Recent advancements in text-to-image (T2I) models enable high-quality image synthesis, yet generating anatomically accurate human figures remains challenging. AI-generated images frequently exhibit distortions such as proliferated limbs, missing fingers, deformed extremities, or fused body parts. Existing evaluation metrics like Inception Score (IS) and Fréchet Inception Distance (FID) lack the granularity to detect these distortions, while human preference-based metrics focus on abstract quality assessments rather than anatomical fidelity. To address this gap, we establish the first standards for identifying human body distortions in AI-generated images and introduce Distortion-5K, a comprehensive dataset comprising 4,700 annotated images of normal and malformed human figures across diverse styles and distortion types. Based on this dataset, we propose ViT-HD, a Vision Transformer-based model tailored for detecting human body distortions in AI-generated images, which outperforms state-of-the-art segmentation models and visual language models, achieving an F1 score of 0.899 and IoU of 0.831 on distortion localization. Additionally, we construct the Human Distortion Benchmark with 500 human-centric prompts to evaluate four popular T2I models using trained ViT-HD, revealing that nearly 50\% of generated images contain distortions. This work pioneers a systematic approach to evaluating anatomical accuracy in AI-generated humans, offering tools to advance the fidelity of T2I models and their real-world applicability. The Distortion-5K dataset, trained ViT-HD will soon be released in our GitHub repository: \href{https://github.com/TheRoadQaQ/Predicting-Distortion}{https://github.com/TheRoadQaQ/Predicting-Distortion}.
format Preprint
id arxiv_https___arxiv_org_abs_2503_00811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating and Predicting Distorted Human Body Parts for Generated Images
Ma, Lu
Cao, Kaibo
Liang, Hao
Lin, Jiaxin
Li, Zhuang
Liu, Yuhong
Zhang, Jihong
Zhang, Wentao
Cui, Bin
Computer Vision and Pattern Recognition
Recent advancements in text-to-image (T2I) models enable high-quality image synthesis, yet generating anatomically accurate human figures remains challenging. AI-generated images frequently exhibit distortions such as proliferated limbs, missing fingers, deformed extremities, or fused body parts. Existing evaluation metrics like Inception Score (IS) and Fréchet Inception Distance (FID) lack the granularity to detect these distortions, while human preference-based metrics focus on abstract quality assessments rather than anatomical fidelity. To address this gap, we establish the first standards for identifying human body distortions in AI-generated images and introduce Distortion-5K, a comprehensive dataset comprising 4,700 annotated images of normal and malformed human figures across diverse styles and distortion types. Based on this dataset, we propose ViT-HD, a Vision Transformer-based model tailored for detecting human body distortions in AI-generated images, which outperforms state-of-the-art segmentation models and visual language models, achieving an F1 score of 0.899 and IoU of 0.831 on distortion localization. Additionally, we construct the Human Distortion Benchmark with 500 human-centric prompts to evaluate four popular T2I models using trained ViT-HD, revealing that nearly 50\% of generated images contain distortions. This work pioneers a systematic approach to evaluating anatomical accuracy in AI-generated humans, offering tools to advance the fidelity of T2I models and their real-world applicability. The Distortion-5K dataset, trained ViT-HD will soon be released in our GitHub repository: \href{https://github.com/TheRoadQaQ/Predicting-Distortion}{https://github.com/TheRoadQaQ/Predicting-Distortion}.
title Evaluating and Predicting Distorted Human Body Parts for Generated Images
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.00811