Physics-Based Benchmarking Metrics for Multimodal Synthetic Images
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Kishor Datta, Kamal, Marufa, Rahman, Md. Mahfuzur, Rahman, Fahad, Haque, Mohd Ariful, Siddique, Sunzida |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLCE: A Knowledge-Enhanced Framework for Image Description in Disaster Assessment
by: Rahman, Md. Mahfuzur, et al.
Published: (2025)
by: Rahman, Md. Mahfuzur, et al.
Published: (2025)
UAV (Unmanned Aerial Vehicles): Diverse Applications of UAV Datasets in Segmentation, Classification, Detection, and Tracking
by: Rahman, Md. Mahfuzur, et al.
Published: (2024)
by: Rahman, Md. Mahfuzur, et al.
Published: (2024)
Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules
by: Gupta, Kishor Datta, et al.
Published: (2025)
by: Gupta, Kishor Datta, et al.
Published: (2025)
Physical Rule-Guided Convolutional Neural Network
by: Gupta, Kishor Datta, et al.
Published: (2024)
by: Gupta, Kishor Datta, et al.
Published: (2024)
SOK: Exploring Hallucinations and Security Risks in AI-Assisted Software Development with Insights for LLM Deployment
by: Haque, Ariful, et al.
Published: (2025)
by: Haque, Ariful, et al.
Published: (2025)
Continuous Monitoring of Large-Scale Generative AI via Deterministic Knowledge Graph Structures
by: Gupta, Kishor Datta, et al.
Published: (2025)
by: Gupta, Kishor Datta, et al.
Published: (2025)
Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM
by: Haque, Mohd Ariful, et al.
Published: (2025)
by: Haque, Mohd Ariful, et al.
Published: (2025)
Vision-Based Localization and LLM-based Navigation for Indoor Environments
by: Rahimi, Keyan, et al.
Published: (2025)
by: Rahimi, Keyan, et al.
Published: (2025)
BanglaMM-Disaster: A Multimodal Transformer-Based Deep Learning Framework for Multiclass Disaster Classification in Bangla
by: Islam, Ariful, et al.
Published: (2025)
by: Islam, Ariful, et al.
Published: (2025)
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
by: Munir, Mustafa, et al.
Published: (2024)
by: Munir, Mustafa, et al.
Published: (2024)
FUSED-Net: Detecting Traffic Signs with Limited Data
by: Rahman, Md. Atiqur, et al.
Published: (2024)
by: Rahman, Md. Atiqur, et al.
Published: (2024)
TinyLLM: Evaluation and Optimization of Small Language Models for Agentic Tasks on Edge Devices
by: Haque, Mohd Ariful, et al.
Published: (2025)
by: Haque, Mohd Ariful, et al.
Published: (2025)
TextDiffuser-RL: Efficient and Robust Text Layout Optimization for High-Fidelity Text-to-Image Synthesis
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
by: Rahman, Kazi Mahathir, et al.
Published: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
by: Rahman, Mizanur, et al.
Published: (2025)
by: Rahman, Mizanur, et al.
Published: (2025)
V-Zen: Efficient GUI Understanding and Precise Grounding With A Novel Multimodal LLM
by: Rahman, Abdur, et al.
Published: (2024)
by: Rahman, Abdur, et al.
Published: (2024)
An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
by: Hassan, Riad, et al.
Published: (2025)
by: Hassan, Riad, et al.
Published: (2025)
SONICS: Synthetic Or Not -- Identifying Counterfeit Songs
by: Rahman, Md Awsafur, et al.
Published: (2024)
by: Rahman, Md Awsafur, et al.
Published: (2024)
AnoFPDM: Anomaly Segmentation with Forward Process of Diffusion Models for Brain MRI
by: Che, Yiming, et al.
Published: (2024)
by: Che, Yiming, et al.
Published: (2024)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
by: Sajib, Rakib Hossain, et al.
Published: (2026)
by: Sajib, Rakib Hossain, et al.
Published: (2026)
Toward Faithful Segmentation Attribution via Benchmarking and Dual-Evidence Fusion
by: Sakib, Abu Noman Md, et al.
Published: (2026)
by: Sakib, Abu Noman Md, et al.
Published: (2026)
Beyond Core and Penumbra: Bi-Temporal Image-Driven Stroke Evolution Analysis
by: Rahman, Md Sazidur, et al.
Published: (2026)
by: Rahman, Md Sazidur, et al.
Published: (2026)
Automated Toll Management System Using RFID and Image Processing
by: Ahmed, Raihan, et al.
Published: (2024)
by: Ahmed, Raihan, et al.
Published: (2024)
Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models
by: Samin, Niamul Hassan, et al.
Published: (2026)
by: Samin, Niamul Hassan, et al.
Published: (2026)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
by: Rahman, Md Ashikur, et al.
Published: (2026)
by: Rahman, Md Ashikur, et al.
Published: (2026)
Teacher-Guided One-Shot Pruning via Context-Aware Knowledge Distillation
by: Alim, Md. Samiul, et al.
Published: (2025)
by: Alim, Md. Samiul, et al.
Published: (2025)
CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
by: Saha, Aranya, et al.
Published: (2025)
by: Saha, Aranya, et al.
Published: (2025)
Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2025)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
by: Rahman, Ben
Published: (2025)
by: Rahman, Ben
Published: (2025)
RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
by: Feng, Ganlin, et al.
Published: (2026)
by: Feng, Ganlin, et al.
Published: (2026)
Celeb-FBI: A Benchmark Dataset on Human Full Body Images and Age, Gender, Height and Weight Estimation using Deep Learning Approach
by: Debnath, Pronay, et al.
Published: (2024)
by: Debnath, Pronay, et al.
Published: (2024)
Entropy-Driven Genetic Optimization for Deep-Feature-Guided Low-Light Image Enhancement
by: Datta, Nirjhor, et al.
Published: (2025)
by: Datta, Nirjhor, et al.
Published: (2025)
The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
by: Azad, Asif, et al.
Published: (2025)
by: Azad, Asif, et al.
Published: (2025)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
by: Tashdeed, Ishmam, et al.
Published: (2025)
by: Tashdeed, Ishmam, et al.
Published: (2025)
FusionEnsemble-Net: An Attention-Based Ensemble of Spatiotemporal Networks for Multimodal Sign Language Recognition
by: Islam, Md. Milon, et al.
Published: (2025)
by: Islam, Md. Milon, et al.
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
by: Rahman, Md Maklachur, et al.
Published: (2026)
by: Rahman, Md Maklachur, et al.
Published: (2026)
Multimodal Connectome Fusion via Cross-Attention for Autism Spectrum Disorder Classification Using Graph Learning
by: Rahman, Ansar, et al.
Published: (2026)
by: Rahman, Ansar, et al.
Published: (2026)
MK-UNet: Multi-kernel Lightweight CNN for Medical Image Segmentation
by: Rahman, Md Mostafijur, et al.
Published: (2025)
by: Rahman, Md Mostafijur, et al.
Published: (2025)
LoMix: Learnable Weighted Multi-Scale Logits Mixing for Medical Image Segmentation
by: Rahman, Md Mostafijur, et al.
Published: (2025)
by: Rahman, Md Mostafijur, et al.
Published: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
by: Chowdhury, Md Abtahi Majeed, et al.
Published: (2025)
Similar Items
-
VLCE: A Knowledge-Enhanced Framework for Image Description in Disaster Assessment
by: Rahman, Md. Mahfuzur, et al.
Published: (2025) -
UAV (Unmanned Aerial Vehicles): Diverse Applications of UAV Datasets in Segmentation, Classification, Detection, and Tracking
by: Rahman, Md. Mahfuzur, et al.
Published: (2024) -
Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules
by: Gupta, Kishor Datta, et al.
Published: (2025) -
Physical Rule-Guided Convolutional Neural Network
by: Gupta, Kishor Datta, et al.
Published: (2024) -
SOK: Exploring Hallucinations and Security Risks in AI-Assisted Software Development with Insights for LLM Deployment
by: Haque, Ariful, et al.
Published: (2025)