Instruction-Evidence Contrastive Dual-Stream Decoding for Grounded Vision-Language Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Bangde, Yashwant Pravinrao, Roy, Debaditya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
di: Biswas, Shristi Das, et al.
Pubblicazione: (2026)
di: Biswas, Shristi Das, et al.
Pubblicazione: (2026)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
di: Ee, Yeo Keat, et al.
Pubblicazione: (2026)
di: Ee, Yeo Keat, et al.
Pubblicazione: (2026)
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
di: Wang, Xintong, et al.
Pubblicazione: (2024)
di: Wang, Xintong, et al.
Pubblicazione: (2024)
Predicting the Next Action by Modeling the Abstract Goal
di: Roy, Debaditya, et al.
Pubblicazione: (2022)
di: Roy, Debaditya, et al.
Pubblicazione: (2022)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2025)
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2025)
Interaction Region Visual Transformer for Egocentric Action Anticipation
di: Roy, Debaditya, et al.
Pubblicazione: (2022)
di: Roy, Debaditya, et al.
Pubblicazione: (2022)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
di: Verma, Dhruv, et al.
Pubblicazione: (2024)
di: Verma, Dhruv, et al.
Pubblicazione: (2024)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2023)
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2023)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
di: Jaiswal, Shantanu, et al.
Pubblicazione: (2024)
di: Jaiswal, Shantanu, et al.
Pubblicazione: (2024)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
di: Glória-Silva, Diogo, et al.
Pubblicazione: (2024)
di: Glória-Silva, Diogo, et al.
Pubblicazione: (2024)
SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
di: Park, Woohyeon, et al.
Pubblicazione: (2025)
di: Park, Woohyeon, et al.
Pubblicazione: (2025)
Cross-Image Contrastive Decoding: Precise, Lossless Suppression of Language Priors in Large Vision-Language Models
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
GLaRE: A Graph-based Landmark Region Embedding Network for Emotion Recognition
di: Maji, Debasis, et al.
Pubblicazione: (2025)
di: Maji, Debasis, et al.
Pubblicazione: (2025)
Infection-Reasoner: A Compact Vision-Language Model for Wound Infection Classification with Evidence-Grounded Clinical Reasoning
di: Busaranuvong, Palawat, et al.
Pubblicazione: (2026)
di: Busaranuvong, Palawat, et al.
Pubblicazione: (2026)
How Reasoning Influences Intersectional Biases in Vision Language Models
di: Desai, Adit, et al.
Pubblicazione: (2025)
di: Desai, Adit, et al.
Pubblicazione: (2025)
Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
di: Zhang, Jialiang, et al.
Pubblicazione: (2026)
di: Zhang, Jialiang, et al.
Pubblicazione: (2026)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
di: Jin, Hyundong, et al.
Pubblicazione: (2025)
Temporal Contrastive Learning for Video Temporal Reasoning in Large Vision-Language Models
di: Souza, Rafael, et al.
Pubblicazione: (2024)
di: Souza, Rafael, et al.
Pubblicazione: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
di: Wang, Chao, et al.
Pubblicazione: (2025)
di: Wang, Chao, et al.
Pubblicazione: (2025)
Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding
di: Mahdavi, Zahra, et al.
Pubblicazione: (2025)
di: Mahdavi, Zahra, et al.
Pubblicazione: (2025)
Dual-Stream Diffusion for World-Model Augmented Vision-Language-Action Model
di: Won, John, et al.
Pubblicazione: (2025)
di: Won, John, et al.
Pubblicazione: (2025)
Can Vision Transformers with ResNet's Global Features Fairly Authenticate Demographic Faces?
di: Sufian, Abu, et al.
Pubblicazione: (2025)
di: Sufian, Abu, et al.
Pubblicazione: (2025)
StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
di: Wang, Han, et al.
Pubblicazione: (2026)
di: Wang, Han, et al.
Pubblicazione: (2026)
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding
di: Wang, Jiaqi, et al.
Pubblicazione: (2024)
di: Wang, Jiaqi, et al.
Pubblicazione: (2024)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
di: Gröpl, Marcel, et al.
Pubblicazione: (2026)
di: Gröpl, Marcel, et al.
Pubblicazione: (2026)
ViTCN: Vision Transformer Contrastive Network For Reasoning
di: Song, Bo, et al.
Pubblicazione: (2024)
di: Song, Bo, et al.
Pubblicazione: (2024)
C3L: Content Correlated Vision-Language Instruction Tuning Data Generation via Contrastive Learning
di: Ma, Ji, et al.
Pubblicazione: (2024)
di: Ma, Ji, et al.
Pubblicazione: (2024)
Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
di: Huang, Yanxiang, et al.
Pubblicazione: (2026)
di: Huang, Yanxiang, et al.
Pubblicazione: (2026)
Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding
di: Sharma, Shivam, et al.
Pubblicazione: (2026)
di: Sharma, Shivam, et al.
Pubblicazione: (2026)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
di: Man, Yunze, et al.
Pubblicazione: (2025)
di: Man, Yunze, et al.
Pubblicazione: (2025)
iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
di: Yao, Manyi, et al.
Pubblicazione: (2025)
di: Yao, Manyi, et al.
Pubblicazione: (2025)
Visual Alignment of Medical Vision-Language Models for Grounded Radiology Report Generation
di: Bose, Sarosij, et al.
Pubblicazione: (2025)
di: Bose, Sarosij, et al.
Pubblicazione: (2025)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models
di: Mahmood, Hazza, et al.
Pubblicazione: (2026)
di: Mahmood, Hazza, et al.
Pubblicazione: (2026)
MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
di: Wu, Xiyang, et al.
Pubblicazione: (2025)
di: Wu, Xiyang, et al.
Pubblicazione: (2025)
Streaming Video Instruction Tuning
di: Xia, Jiaer, et al.
Pubblicazione: (2025)
di: Xia, Jiaer, et al.
Pubblicazione: (2025)
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
di: Fang, Hao, et al.
Pubblicazione: (2025)
di: Fang, Hao, et al.
Pubblicazione: (2025)
Context-Aware Pesticide Recommendation via Few-Shot Pest Recognition for Precision Agriculture
di: Ghosh, Anirudha, et al.
Pubblicazione: (2026)
di: Ghosh, Anirudha, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
di: Biswas, Shristi Das, et al.
Pubblicazione: (2026) -
Improving Temporal Action Segmentation via Constraint-Aware Decoding
di: Ee, Yeo Keat, et al.
Pubblicazione: (2026) -
Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
di: Wang, Xintong, et al.
Pubblicazione: (2024) -
Predicting the Next Action by Modeling the Abstract Goal
di: Roy, Debaditya, et al.
Pubblicazione: (2022) -
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
di: Rajendiran, Ramanathan, et al.
Pubblicazione: (2025)