Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arif, Nokimul Hasan, Rabby, Shadman, Papon, Md Hefzul Hossain, Ahmed, Sabbir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Experimental Comparison of Light-Weight and Deep CNN Models Across Diverse Datasets
von: Papon, Md. Hefzul Hossain, et al.
Veröffentlicht: (2026)
von: Papon, Md. Hefzul Hossain, et al.
Veröffentlicht: (2026)
Moral Sycophancy in Vision Language Models
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
von: Rabby, Shadman, et al.
Veröffentlicht: (2026)
DExNet: Combining Observations of Domain Adapted Critics for Leaf Disease Classification with Limited Data
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025)
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025)
Prompt2SegCXR:Prompt to Segment All Organs and Diseases in Chest X-rays
von: Zami, Abduz, et al.
Veröffentlicht: (2025)
von: Zami, Abduz, et al.
Veröffentlicht: (2025)
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
Jellyfish Species Identification: A CNN Based Artificial Neural Network Approach
von: Hossen, Md. Sabbir, et al.
Veröffentlicht: (2025)
von: Hossen, Md. Sabbir, et al.
Veröffentlicht: (2025)
BdSLW60: A Word-Level Bangla Sign Language Dataset
von: Rubaiyeat, Husne Ara, et al.
Veröffentlicht: (2024)
von: Rubaiyeat, Husne Ara, et al.
Veröffentlicht: (2024)
Two Decades of Bengali Handwritten Digit Recognition: A Survey
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2022)
von: Rahman, A. B. M. Ashikur, et al.
Veröffentlicht: (2022)
Beyond Symbolic Solving: Multi Chain-of-Thought Voting for Geometric Reasoning in Large Language Models
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
von: Siddique, Md. Abu Bakor, et al.
Veröffentlicht: (2026)
Hallucination of Multimodal Large Language Models: A Survey
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
von: Bai, Zechen, et al.
Veröffentlicht: (2024)
Vision-Language Models for Automated Chest X-ray Interpretation: Leveraging ViT and GPT-2
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
von: Islam, Md. Rakibul, et al.
Veröffentlicht: (2025)
Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks
von: Shawon, Jubayer Ahmed Bhuiyan, et al.
Veröffentlicht: (2025)
von: Shawon, Jubayer Ahmed Bhuiyan, et al.
Veröffentlicht: (2025)
Robustness of Vision Language Models Against Split-Image Harmful Input Attacks
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Rafi Ur, et al.
Veröffentlicht: (2026)
LISA: A Layer-wise Integration and Suppression Approach for Hallucination Mitigation in Multimodal Large Language Models
von: Guo, Zhihui, et al.
Veröffentlicht: (2025)
von: Guo, Zhihui, et al.
Veröffentlicht: (2025)
FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs
von: Yan, Bowen, et al.
Veröffentlicht: (2024)
von: Yan, Bowen, et al.
Veröffentlicht: (2024)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
MangoLeafViT: Leveraging Lightweight Vision Transformer with Runtime Augmentation for Efficient Mango Leaf Disease Classification
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
von: Chowdhury, Rafi Hassan, et al.
Veröffentlicht: (2025)
A Computer Vision Based Approach for Stalking Detection Using a CNN-LSTM-MLP Hybrid Fusion Model
von: Hasan, Murad, et al.
Veröffentlicht: (2024)
von: Hasan, Murad, et al.
Veröffentlicht: (2024)
Thinking Like a Botanist: Challenging Multimodal Language Models with Intent-Driven Chain-of-Inquiry
von: Sakib, Syed Nazmus, et al.
Veröffentlicht: (2026)
von: Sakib, Syed Nazmus, et al.
Veröffentlicht: (2026)
FUSED-Net: Detecting Traffic Signs with Limited Data
von: Rahman, Md. Atiqur, et al.
Veröffentlicht: (2024)
von: Rahman, Md. Atiqur, et al.
Veröffentlicht: (2024)
Efficient Preemptive Robustification with Image Sharpening
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
Score-Control for Hallucination Reduction in Diffusion Models
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
Fine-Tuned CNN-Based Approach for Multi-Class Mango Leaf Disease Detection
von: Ahmmed, Jalal, et al.
Veröffentlicht: (2025)
von: Ahmmed, Jalal, et al.
Veröffentlicht: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
von: He, Zhentao, et al.
Veröffentlicht: (2025)
von: He, Zhentao, et al.
Veröffentlicht: (2025)
Performance Analysis of Few-Shot Learning Approaches for Bangla Handwritten Character and Digit Recognition
von: Ahamed, Mehedi, et al.
Veröffentlicht: (2025)
von: Ahamed, Mehedi, et al.
Veröffentlicht: (2025)
Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
von: Hossain, Md Zarif, et al.
Veröffentlicht: (2024)
Counting Through Occlusion: Framework for Open World Amodal Counting
von: Arib, Safaeid Hossain, et al.
Veröffentlicht: (2025)
von: Arib, Safaeid Hossain, et al.
Veröffentlicht: (2025)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
von: Ding, Peng, et al.
Veröffentlicht: (2024)
von: Ding, Peng, et al.
Veröffentlicht: (2024)
Few-Shot Learning Pipeline for Monkeypox Skin Disease Classification Using CNN Feature Extractors
von: Rashid, Md. Safirur, et al.
Veröffentlicht: (2026)
von: Rashid, Md. Safirur, et al.
Veröffentlicht: (2026)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
von: Li, Jiale, et al.
Veröffentlicht: (2025)
von: Li, Jiale, et al.
Veröffentlicht: (2025)
BeyondPixels: A Comprehensive Review of the Evolution of Neural Radiance Fields
von: Rabby, AKM Shahariar Azad, et al.
Veröffentlicht: (2023)
von: Rabby, AKM Shahariar Azad, et al.
Veröffentlicht: (2023)
DL$^3$M: A Vision-to-Language Framework for Expert-Level Medical Reasoning through Deep Learning and Large Language Models
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
von: Hasan, Md. Najib, et al.
Veröffentlicht: (2025)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
ODE: Open-Set Evaluation of Hallucinations in Multimodal Large Language Models
von: Tu, Yahan, et al.
Veröffentlicht: (2024)
von: Tu, Yahan, et al.
Veröffentlicht: (2024)
OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis
von: Ahmed, Istiak, et al.
Veröffentlicht: (2025)
von: Ahmed, Istiak, et al.
Veröffentlicht: (2025)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
von: Li, Chaoyu, et al.
Veröffentlicht: (2024)
von: Li, Chaoyu, et al.
Veröffentlicht: (2024)
An Embedded Real-time Object Alert System for Visually Impaired: A Monocular Depth Estimation based Approach through Computer Vision
von: Anjom, Jareen, et al.
Veröffentlicht: (2025)
von: Anjom, Jareen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Experimental Comparison of Light-Weight and Deep CNN Models Across Diverse Datasets
von: Papon, Md. Hefzul Hossain, et al.
Veröffentlicht: (2026) -
Moral Sycophancy in Vision Language Models
von: Rabby, Shadman, et al.
Veröffentlicht: (2026) -
DExNet: Combining Observations of Domain Adapted Critics for Leaf Disease Classification with Limited Data
von: Ahmed, Sabbir, et al.
Veröffentlicht: (2025) -
PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model
von: Arif, Kazi Hasan Ibn, et al.
Veröffentlicht: (2025) -
Prompt2SegCXR:Prompt to Segment All Organs and Diseases in Chest X-rays
von: Zami, Abduz, et al.
Veröffentlicht: (2025)