Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866918488750161920 |
|---|---|
| author | Gupta, Kishor Datta Kamal, Marufa Rahman, Md. Mahfuzur Rahman, Fahad Haque, Mohd Ariful Siddique, Sunzida |
| author_facet | Gupta, Kishor Datta Kamal, Marufa Rahman, Md. Mahfuzur Rahman, Fahad Haque, Mohd Ariful Siddique, Sunzida |
| contents | Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes a Physics-Constrained Multimodal Data Evaluation (PCMDE) metric combining large language models with reasoning, knowledge based mapping and vision-language models to overcome these limitations. The architecture is comprised of three main stages: (1) feature extraction of spatial and semantic information with multimodal features through object detection and VLMs; (2) Confidence-Weighted Component Fusion for adaptive component-level validation; and (3) physics-guided reasoning using large language models for structural and relational constraints (e.g., alignment, position, consistency) enforcement. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2511_15204 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Physics-Based Benchmarking Metrics for Multimodal Synthetic Images Gupta, Kishor Datta Kamal, Marufa Rahman, Md. Mahfuzur Rahman, Fahad Haque, Mohd Ariful Siddique, Sunzida Computer Vision and Pattern Recognition Artificial Intelligence Current state of the art measures like BLEU, CIDEr, VQA score, SigLIP-2 and CLIPScore are often unable to capture semantic or structural accuracy, especially for domain-specific or context-dependent scenarios. For this, this paper proposes a Physics-Constrained Multimodal Data Evaluation (PCMDE) metric combining large language models with reasoning, knowledge based mapping and vision-language models to overcome these limitations. The architecture is comprised of three main stages: (1) feature extraction of spatial and semantic information with multimodal features through object detection and VLMs; (2) Confidence-Weighted Component Fusion for adaptive component-level validation; and (3) physics-guided reasoning using large language models for structural and relational constraints (e.g., alignment, position, consistency) enforcement. |
| title | Physics-Based Benchmarking Metrics for Multimodal Synthetic Images |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2511.15204 |