Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Deria, Ankan, Dukre, Adinath Madhavrao, Tang, Feilong, Atito, Sara, Roy, Sudipta, Awais, Muhammad, Khan, Muhammad Haris, Razzak, Imran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
Robust Atypical Mitosis Classification with DenseNet121: Stain-Aware Augmentation and Hybrid Loss for Domain Generalization
by: Dukre, Adinath, et al.
Published: (2025)
by: Dukre, Adinath, et al.
Published: (2025)
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment
by: Azeez, Mohammad Anas, et al.
Published: (2026)
by: Azeez, Mohammad Anas, et al.
Published: (2026)
TuLaBM: Tumor-Biased Latent Bridge Matching for Contrast-Enhanced MRI Synthesis
by: Rege, Atharva, et al.
Published: (2026)
by: Rege, Atharva, et al.
Published: (2026)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
by: Deria, Ankan, et al.
Published: (2026)
by: Deria, Ankan, et al.
Published: (2026)
Domain Adaptation Without the Compute Burden for Efficient Whole Slide Image Analysis
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
Information theoretic underpinning of self-supervised learning by clustering
by: Kittler, Josef, et al.
Published: (2026)
by: Kittler, Josef, et al.
Published: (2026)
PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
by: Alex, Tony, et al.
Published: (2025)
by: Alex, Tony, et al.
Published: (2025)
MuGa-VTON: Multi-Garment Virtual Try-On via Diffusion Transformers with Prompt Customization
by: Deria, Ankan, et al.
Published: (2025)
by: Deria, Ankan, et al.
Published: (2025)
C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging
by: Marikkar, Umar, et al.
Published: (2025)
by: Marikkar, Umar, et al.
Published: (2025)
Channel-Aware Probing for Multi-Channel Imaging
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
DC-ViT: Modulating Spatial and Channel Interactions for Multi-Channel Images
by: Marikkar, Umar, et al.
Published: (2026)
by: Marikkar, Umar, et al.
Published: (2026)
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training
by: Kumar, Komal, et al.
Published: (2026)
by: Kumar, Komal, et al.
Published: (2026)
DeepChest: Dynamic Gradient-Free Task Weighting for Effective Multi-Task Learning in Chest X-ray Classification
by: Mohamed, Youssef, et al.
Published: (2025)
by: Mohamed, Youssef, et al.
Published: (2025)
GeoVLM: Improving Automated Vehicle Geolocalisation Using Vision-Language Matching
by: Dagda, Barkin, et al.
Published: (2025)
by: Dagda, Barkin, et al.
Published: (2025)
Key-Conditioned Orthonormal Transform Gating (K-OTG): Multi-Key Access Control with Hidden-State Scrambling for LoRA-Tuned Models
by: Khan, Muhammad Haris
Published: (2025)
by: Khan, Muhammad Haris
Published: (2025)
Is Monotonic Sampling Necessary in Diffusion Models?
by: Khan, Muhammad Haris
Published: (2026)
by: Khan, Muhammad Haris
Published: (2026)
SafeBench-Seq: A Homology-Clustered, CPU-Only Baseline for Protein Hazard Screening with Physicochemical/Composition Features and Cluster-Aware Confidence Intervals
by: Khan, Muhammad Haris
Published: (2025)
by: Khan, Muhammad Haris
Published: (2025)
Investigating Self-Supervised Methods for Label-Efficient Learning
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
by: Atito, Sara, et al.
Published: (2022)
by: Atito, Sara, et al.
Published: (2022)
Pseudo Labelling for Enhanced Masked Autoencoders
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
by: Nandam, Srinivasa Rao, et al.
Published: (2024)
Few-Label Multimodal Modeling of SNP Variants and ECG Phenotypes Using Large Language Models for Cardiovascular Risk Stratification
by: Menon, Niranjana Arun, et al.
Published: (2025)
by: Menon, Niranjana Arun, et al.
Published: (2025)
DeLo: Dual Decomposed Low-Rank Experts Collaboration for Continual Missing Modality Learning
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
Rethinking Positive Pairs in Contrastive Learning
by: Wu, Jiantao, et al.
Published: (2024)
by: Wu, Jiantao, et al.
Published: (2024)
DailyMAE: Towards Pretraining Masked Autoencoders in One Day
by: Wu, Jiantao, et al.
Published: (2024)
by: Wu, Jiantao, et al.
Published: (2024)
Advancements in deep learning for Alzheimer's disease diagnosis: A comprehensive exploration and critical analysis of neuroimaging approaches
by: Fakhri Alam Khan, et al.
Published: (2024)
by: Fakhri Alam Khan, et al.
Published: (2024)
Robust and Label-Efficient Deep Waste Detection
by: Abid, Hassan, et al.
Published: (2025)
by: Abid, Hassan, et al.
Published: (2025)
CountZES: Counting via Zero-Shot Exemplar Selection
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
by: Khan, Muhammad Haris, et al.
Published: (2025)
by: Khan, Muhammad Haris, et al.
Published: (2025)
Probabilistically Aligned View-unaligned Clustering with Adaptive Template Selection
by: Dong, Wenhua, et al.
Published: (2024)
by: Dong, Wenhua, et al.
Published: (2024)
How Effectively Can Large Language Models Connect SNP Variants and ECG Phenotypes for Cardiovascular Risk Prediction?
by: Menon, Niranjana Arun, et al.
Published: (2025)
by: Menon, Niranjana Arun, et al.
Published: (2025)
TESL-Net: A Transformer-Enhanced CNN for Accurate Skin Lesion Segmentation
by: Iqbal, Shahzaib, et al.
Published: (2024)
by: Iqbal, Shahzaib, et al.
Published: (2024)
Pose-Guided Self-Training with Two-Stage Clustering for Unsupervised Landmark Discovery
by: Tourani, Siddharth, et al.
Published: (2024)
by: Tourani, Siddharth, et al.
Published: (2024)
LATA: Laplacian-Assisted Transductive Adaptation for Conformal Uncertainty in Medical VLMs
by: Bozorgtabar, Behzad, et al.
Published: (2026)
by: Bozorgtabar, Behzad, et al.
Published: (2026)
BioVLM: Routing Prompts, Not Parameters, for Cross-Modality Generalization in Biomedical VLMs
by: Singha, Mainak, et al.
Published: (2026)
by: Singha, Mainak, et al.
Published: (2026)
Divergent Domains, Convergent Grading: Enhancing Generalization in Diabetic Retinopathy Grading
by: Chokuwa, Sharon, et al.
Published: (2024)
by: Chokuwa, Sharon, et al.
Published: (2024)
Noise-Tolerant Few-Shot Unsupervised Adapter for Vision-Language Models
by: Ali, Eman, et al.
Published: (2023)
by: Ali, Eman, et al.
Published: (2023)
StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
by: Yang, Haolin, et al.
Published: (2025)
by: Yang, Haolin, et al.
Published: (2025)
On Deep Learning for computing the Dynamic Initial Margin and Margin Value Adjustment
by: Villarino, Joel P., et al.
Published: (2024)
by: Villarino, Joel P., et al.
Published: (2024)
Similar Items
-
MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
by: Deria, Ankan, et al.
Published: (2026) -
Robust Atypical Mitosis Classification with DenseNet121: Stain-Aware Augmentation and Hybrid Loss for Domain Generalization
by: Dukre, Adinath, et al.
Published: (2025) -
See Fair, Speak Truth: Equitable Attention Improves Grounding and Reduces Hallucination in Vision-Language Alignment
by: Azeez, Mohammad Anas, et al.
Published: (2026) -
TuLaBM: Tumor-Biased Latent Bridge Matching for Contrast-Enhanced MRI Synthesis
by: Rege, Atharva, et al.
Published: (2026) -
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)