IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Islam, Md Touhidul, Kabir, Imran, Reza, Md Alimoor, Billah, Syed Masum |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
von: Kabir, Imran, et al.
Veröffentlicht: (2025)
Wheeler: A Three-Wheeled Input Device for Usable, Efficient, and Versatile Non-Visual Interaction
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
Demonstration of Wheeler: A Three-Wheeled Input Device for Usable, Efficient, and Versatile Non-Visual Interaction
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)
Rice Leaf Disease Detection: A Comparative Study Between CNN, Transformer and Non-neural Network Architectures
von: Mehnaz, Samia, et al.
Veröffentlicht: (2025)
von: Mehnaz, Samia, et al.
Veröffentlicht: (2025)
Vision-Based Lane Following and Traffic Sign Recognition for Resource-Constrained Autonomous Vehicles
von: Islam, Md Tanjemul, et al.
Veröffentlicht: (2026)
von: Islam, Md Tanjemul, et al.
Veröffentlicht: (2026)
Shaping Credibility Judgments in Human-GenAI Partnership via Weaker LLMs: A Transactive Memory Perspective on AI Literacy
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
HeBA: Heterogeneous Bottleneck Adapters for Robust Vision-Language Models
von: Islam, Md Jahidul
Veröffentlicht: (2026)
von: Islam, Md Jahidul
Veröffentlicht: (2026)
Quantitative Currency Evaluation in Low-Resource Settings through Pattern Analysis to Assist Visually Impaired Users
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam, et al.
Veröffentlicht: (2025)
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
von: Rahman, Md Ashikur, et al.
Veröffentlicht: (2026)
Automated Wicket-Taking Delivery Segmentation and Trajectory-Based Dismissal-Zone Analysis in Cricket Videos Using OCR-Guided YOLOv8
von: Karmoker, Joy, et al.
Veröffentlicht: (2025)
von: Karmoker, Joy, et al.
Veröffentlicht: (2025)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Reliable Deep Learning for Small-Scale Classifications: Experiments on Real-World Image Datasets from Bangladesh
von: Suny, Alfe, et al.
Veröffentlicht: (2026)
von: Suny, Alfe, et al.
Veröffentlicht: (2026)
Blending 3D Geometry and Machine Learning for Multi-View Stereopsis
von: Vats, Vibhas, et al.
Veröffentlicht: (2025)
von: Vats, Vibhas, et al.
Veröffentlicht: (2025)
A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answering
von: Hossain, Md. Zahid, et al.
Veröffentlicht: (2026)
von: Hossain, Md. Zahid, et al.
Veröffentlicht: (2026)
A Computer Vision Based Approach for Stalking Detection Using a CNN-LSTM-MLP Hybrid Fusion Model
von: Hasan, Murad, et al.
Veröffentlicht: (2024)
von: Hasan, Murad, et al.
Veröffentlicht: (2024)
In-Depth Analysis of Automated Acne Disease Recognition and Classification
von: Jeny, Afsana Ahsan, et al.
Veröffentlicht: (2025)
von: Jeny, Afsana Ahsan, et al.
Veröffentlicht: (2025)
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
von: Kabir, Raihan, et al.
Veröffentlicht: (2024)
ReHARK: Refined Hybrid Adaptive RBF Kernels for Robust One-Shot Vision-Language Adaptation
von: Islam, Md Jahidul
Veröffentlicht: (2026)
von: Islam, Md Jahidul
Veröffentlicht: (2026)
ChartZero: Synthetic Priors Enable Zero Shot Chart Data Extraction
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2026)
An Explainable Vision-Language Model Framework with Adaptive PID-Tversky Loss for Lumbar Spinal Stenosis Diagnosis
von: Sk., Md. Sajeebul Islam, et al.
Veröffentlicht: (2026)
von: Sk., Md. Sajeebul Islam, et al.
Veröffentlicht: (2026)
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment
von: Azeez, Mohammad Anas, et al.
Veröffentlicht: (2026)
von: Azeez, Mohammad Anas, et al.
Veröffentlicht: (2026)
GC-MVSNet: Multi-View, Multi-Scale, Geometrically-Consistent Multi-View Stereo
von: Vats, Vibhas K., et al.
Veröffentlicht: (2023)
von: Vats, Vibhas K., et al.
Veröffentlicht: (2023)
Comparative Performance Analysis of Transformer-Based Pre-Trained Models for Detecting Keratoconus Disease
von: Ahmed, Nayeem, et al.
Veröffentlicht: (2024)
von: Ahmed, Nayeem, et al.
Veröffentlicht: (2024)
When Language Model Guides Vision: Grounding DINO for Cattle Muzzle Detection
von: Dulal, Rabin, et al.
Veröffentlicht: (2025)
von: Dulal, Rabin, et al.
Veröffentlicht: (2025)
Beyond Dominant Patches: Spatial Credit Redistribution For Grounded Vision-Language Models
von: Samin, Niamul Hassan, et al.
Veröffentlicht: (2026)
von: Samin, Niamul Hassan, et al.
Veröffentlicht: (2026)
Leveraging Pre-trained CNNs for Efficient Feature Extraction in Rice Leaf Disease Classification
von: Sobuj, Md. Shohanur Islam, et al.
Veröffentlicht: (2024)
von: Sobuj, Md. Shohanur Islam, et al.
Veröffentlicht: (2024)
ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time Interaction
von: Islam, Md Zabirul, et al.
Veröffentlicht: (2025)
von: Islam, Md Zabirul, et al.
Veröffentlicht: (2025)
Learning Visual Grounding from Generative Vision and Language Model
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
von: Wang, Shijie, et al.
Veröffentlicht: (2024)
VFM-VLM: Vision Foundation Model and Vision Language Model based Visual Comparison for 3D Pose Estimation
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
von: Sarowar, Md Selim, et al.
Veröffentlicht: (2025)
A Domain-Adapted Lightweight Ensemble for Resource-Efficient Few-Shot Plant Disease Classification
von: Islam, Anika, et al.
Veröffentlicht: (2025)
von: Islam, Anika, et al.
Veröffentlicht: (2025)
ELMF4EggQ: Ensemble Learning with Multimodal Feature Fusion for Non-Destructive Egg Quality Assessment
von: Hassan, Md Zahim, et al.
Veröffentlicht: (2025)
von: Hassan, Md Zahim, et al.
Veröffentlicht: (2025)
W-DUALMINE: Reliability-Weighted Dual-Expert Fusion With Residual Correlation Preservation for Medical Image Fusion
von: Islam, Md. Jahidul
Veröffentlicht: (2026)
von: Islam, Md. Jahidul
Veröffentlicht: (2026)
An Image Dataset of Common Skin Diseases of Bangladesh and Benchmarking Performance with Machine Learning Models
von: Hossain, Sazzad, et al.
Veröffentlicht: (2026)
von: Hossain, Sazzad, et al.
Veröffentlicht: (2026)
CS-Mixer: A Cross-Scale Vision MLP Model with Spatial-Channel Mixing
von: Cui, Jonathan, et al.
Veröffentlicht: (2023)
von: Cui, Jonathan, et al.
Veröffentlicht: (2023)
PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
von: Chen, Minbing, et al.
Veröffentlicht: (2026)
Contextual Checkerboard Denoise -- A Novel Neural Network-Based Approach for Classification-Aware OCT Image Denoising
von: Islam, Md. Touhidul, et al.
Veröffentlicht: (2024)
von: Islam, Md. Touhidul, et al.
Veröffentlicht: (2024)
ElderFallGuard: Real-Time IoT and Computer Vision-Based Fall Detection System for Elderly Safety
von: Riahi, Tasrifur, et al.
Veröffentlicht: (2025)
von: Riahi, Tasrifur, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Identifying Crucial Objects in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024) -
A Dataset for Crucial Object Recognition in Blind and Low-Vision Individuals' Navigation
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024) -
Logic-RAG: Augmenting Large Multimodal Models with Visual-Spatial Knowledge for Road Scene Understanding
von: Kabir, Imran, et al.
Veröffentlicht: (2025) -
Wheeler: A Three-Wheeled Input Device for Usable, Efficient, and Versatile Non-Visual Interaction
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024) -
Demonstration of Wheeler: A Three-Wheeled Input Device for Usable, Efficient, and Versatile Non-Visual Interaction
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2024)