Physics-Based Benchmarking Metrics for Multimodal Synthetic Images

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gupta, Kishor Datta, Kamal, Marufa, Rahman, Md. Mahfuzur, Rahman, Fahad, Haque, Mohd Ariful, Siddique, Sunzida
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918488750161920
author Gupta, Kishor Datta
Kamal, Marufa
Rahman, Md. Mahfuzur
Rahman, Fahad
Haque, Mohd Ariful
Siddique, Sunzida
author_facet Gupta, Kishor Datta
Kamal, Marufa
Rahman, Md. Mahfuzur
Rahman, Fahad
Haque, Mohd Ariful
Siddique, Sunzida
contents Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes a Physics-Constrained Multimodal Data Evaluation (PCMDE) metric combining large language models with reasoning, knowledge based mapping and vision-language models to overcome these limitations. The architecture is comprised of three main stages: (1) feature extraction of spatial and semantic information with multimodal features through object detection and VLMs; (2) Confidence-Weighted Component Fusion for adaptive component-level validation; and (3) physics-guided reasoning using large language models for structural and relational constraints (e.g., alignment, position, consistency) enforcement.
format Preprint
id arxiv_https___arxiv_org_abs_2511_15204
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
Gupta, Kishor Datta
Kamal, Marufa
Rahman, Md. Mahfuzur
Rahman, Fahad
Haque, Mohd Ariful
Siddique, Sunzida
Computer Vision and Pattern Recognition
Artificial Intelligence
Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes a Physics-Constrained Multimodal Data Evaluation (PCMDE) metric combining large language models with reasoning, knowledge based mapping and vision-language models to overcome these limitations. The architecture is comprised of three main stages: (1) feature extraction of spatial and semantic information with multimodal features through object detection and VLMs; (2) Confidence-Weighted Component Fusion for adaptive component-level validation; and (3) physics-guided reasoning using large language models for structural and relational constraints (e.g., alignment, position, consistency) enforcement.
title Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2511.15204