The Art of Saying "Maybe": A Conformal Lens for Uncertainty Benchmarking in VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Azad, Asif, Hossain, Mohammad Sadat, Shanto, MD Sadik Hossain, Rahman, M Saifur, Parvez, Md Rizwan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
by: Rashid, Md Rafi Ur, et al.
Published: (2026)
by: Rashid, Md Rafi Ur, et al.
Published: (2026)
AttMetNet: Attention-Enhanced Deep Neural Network for Methane Plume Detection in Sentinel-2 Satellite Imagery
by: Ahsan, Rakib, et al.
Published: (2025)
by: Ahsan, Rakib, et al.
Published: (2025)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
by: Paul, Dhiman, et al.
Published: (2024)
by: Paul, Dhiman, et al.
Published: (2024)
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
by: Sadik, Md Nahid, et al.
Published: (2024)
by: Sadik, Md Nahid, et al.
Published: (2024)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
by: Tashdeed, Ishmam, et al.
Published: (2025)
by: Tashdeed, Ishmam, et al.
Published: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
by: Anonto, Riad Ahmed, et al.
Published: (2025)
by: Anonto, Riad Ahmed, et al.
Published: (2025)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
Integrating Skeleton Based Representations for Robust Yoga Pose Classification Using Deep Learning Models
by: Mohiuddin, Mohammed, et al.
Published: (2025)
by: Mohiuddin, Mohammed, et al.
Published: (2025)
DFCon: Attention-Driven Supervised Contrastive Learning for Robust Deepfake Detection
by: Shanto, MD Sadik Hossain, et al.
Published: (2025)
by: Shanto, MD Sadik Hossain, et al.
Published: (2025)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026)
by: Bhaskar, Paramananda, et al.
Published: (2026)
VLAgeBench: Benchmarking Large Vision-Language Models for Zero-Shot Human Age Estimation
by: Sajib, Rakib Hossain, et al.
Published: (2026)
by: Sajib, Rakib Hossain, et al.
Published: (2026)
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment
by: Azad, Asif, et al.
Published: (2026)
by: Azad, Asif, et al.
Published: (2026)
Reliable Deep Learning for Small-Scale Classifications: Experiments on Real-World Image Datasets from Bangladesh
by: Suny, Alfe, et al.
Published: (2026)
by: Suny, Alfe, et al.
Published: (2026)
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
by: Zhang, Yuyou, et al.
Published: (2025)
by: Zhang, Yuyou, et al.
Published: (2025)
Toward Reliable and Explainable Nail Disease Classification: Leveraging Adversarial Training and Grad-CAM Visualization
by: Hossain, Farzia, et al.
Published: (2026)
by: Hossain, Farzia, et al.
Published: (2026)
Multi-Level Bidirectional Decoder Interaction for Uncertainty-Aware Breast Ultrasound Analysis
by: Shafi, Abdullah Al, et al.
Published: (2026)
by: Shafi, Abdullah Al, et al.
Published: (2026)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
by: Tonmoy, Moshiur Rahman, et al.
Published: (2025)
CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
by: Hossain, Md Tanvir, et al.
Published: (2025)
by: Hossain, Md Tanvir, et al.
Published: (2025)
Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
by: Peter, Ojonugwa Oluwafemi Ejiga, et al.
Published: (2026)
An Efficient Dual-Line Decoder Network with Multi-Scale Convolutional Attention for Multi-organ Segmentation
by: Hassan, Riad, et al.
Published: (2025)
by: Hassan, Riad, et al.
Published: (2025)
Evaluating YOLO Architectures: Implications for Real-Time Vehicle Detection in Urban Environments of Bangladesh
by: Hossain, Ha Meem, et al.
Published: (2025)
by: Hossain, Ha Meem, et al.
Published: (2025)
MF-GCN: A Multi-Frequency Graph Convolutional Network for Tri-Modal Depression Detection Using Eye-Tracking, Facial, and Acoustic Features
by: Rahman, Sejuti, et al.
Published: (2025)
by: Rahman, Sejuti, et al.
Published: (2025)
Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
by: Hossain, Shamima
Published: (2025)
by: Hossain, Shamima
Published: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
by: Burgess, James, et al.
Published: (2026)
by: Burgess, James, et al.
Published: (2026)
Bangladeshi Street Food Calorie Estimation Using Improved YOLOv8 and Regression Model
by: Dhar, Aparup, et al.
Published: (2025)
by: Dhar, Aparup, et al.
Published: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
by: Hossain, Md Zarif, et al.
Published: (2024)
by: Hossain, Md Zarif, et al.
Published: (2024)
KDC-Diff: A Latent-Aware Diffusion Model with Knowledge Retention for Memory-Efficient Image Generation
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
Probabilistic Feature Imputation and Uncertainty-Aware Multimodal Federated Aggregation
by: Shahid, Nafis Fuad, et al.
Published: (2026)
by: Shahid, Nafis Fuad, et al.
Published: (2026)
HybridSolarNet: A Lightweight and Explainable EfficientNet-CBAM Architecture for Real-Time Solar Panel Fault Detection
by: Hossain, Md. Asif, et al.
Published: (2026)
by: Hossain, Md. Asif, et al.
Published: (2026)
From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities
by: Ishmam, Md Farhan, et al.
Published: (2023)
by: Ishmam, Md Farhan, et al.
Published: (2023)
Beyond Fixed Thresholds and Domain-Specific Benchmarks for Explainable Multi-Task Classification in Autonomous Vehicles
by: Azad, Maryam Sadat Hosseini, et al.
Published: (2026)
by: Azad, Maryam Sadat Hosseini, et al.
Published: (2026)
LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
by: Bozorgtabar, Behzad, et al.
Published: (2026)
by: Bozorgtabar, Behzad, et al.
Published: (2026)
SRLoRA: Subspace Recomposition in Low-Rank Adaptation via Importance-Based Fusion and Reinitialization
by: Yang, Haodong, et al.
Published: (2025)
by: Yang, Haodong, et al.
Published: (2025)
ForCM: Forest Cover Mapping from Multispectral Sentinel-2 Image by Integrating Deep Learning with Object-Based Image Analysis
by: Haque, Maisha, et al.
Published: (2025)
by: Haque, Maisha, et al.
Published: (2025)
Autonomous Navigation of Cloud-Controlled Quadcopters in Confined Spaces Using Multi-Modal Perception and LLM-Driven High Semantic Reasoning
by: Ahmmad, Shoaib, et al.
Published: (2025)
by: Ahmmad, Shoaib, et al.
Published: (2025)
PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
by: Sakib, Syed Nazmus, et al.
Published: (2025)
by: Sakib, Syed Nazmus, et al.
Published: (2025)
InfiltrNet: Dual-Branch CNN-Transformer Architecture for Brain Tumor Infiltration Risk Prediction
by: Hossain, S M Asif, et al.
Published: (2026)
by: Hossain, S M Asif, et al.
Published: (2026)
Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
by: Hossain, Md. Iqbal, et al.
Published: (2025)
by: Hossain, Md. Iqbal, et al.
Published: (2025)
An empirical study for the early detection of Mpox from skin lesion images using pretrained CNN models leveraging XAI technique
by: Rahim, Mohammad Asifur, et al.
Published: (2025)
by: Rahim, Mohammad Asifur, et al.
Published: (2025)
Similar Items
-
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
by: Rashid, Md Rafi Ur, et al.
Published: (2026) -
AttMetNet: Attention-Enhanced Deep Neural Network for Methane Plume Detection in Sentinel-2 Satellite Imagery
by: Ahsan, Rakib, et al.
Published: (2025) -
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval
by: Paul, Dhiman, et al.
Published: (2024) -
Real-Time Detection and Analysis of Vehicles and Pedestrians using Deep Learning
by: Sadik, Md Nahid, et al.
Published: (2024) -
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
by: Tashdeed, Ishmam, et al.
Published: (2025)