GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zeshang, Zhang, Shuoyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
von: Maeda, Koki, et al.
Veröffentlicht: (2024)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
von: Dong, Hao, et al.
Veröffentlicht: (2025)
von: Dong, Hao, et al.
Veröffentlicht: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
von: Qin, Guangshuo, et al.
Veröffentlicht: (2026)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
MAMS: Model-Agnostic Module Selection Framework for Video Captioning
von: Lee, Sangho, et al.
Veröffentlicht: (2025)
von: Lee, Sangho, et al.
Veröffentlicht: (2025)
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
von: Lavoie, Samuel, et al.
Veröffentlicht: (2024)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
CaptionFool: Universal Image Captioning Model Attacks
von: Parekh, Swapnil
Veröffentlicht: (2026)
von: Parekh, Swapnil
Veröffentlicht: (2026)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
SAUCE: Selective Concept Unlearning in Vision-Language Models with Sparse Autoencoders
von: Li, Qing, et al.
Veröffentlicht: (2025)
von: Li, Qing, et al.
Veröffentlicht: (2025)
RS-MoE: A Vision-Language Model with Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering
von: Lin, Hui, et al.
Veröffentlicht: (2024)
von: Lin, Hui, et al.
Veröffentlicht: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
von: Yu, Bo, et al.
Veröffentlicht: (2026)
von: Yu, Bo, et al.
Veröffentlicht: (2026)
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
von: Zhang, Jianke, et al.
Veröffentlicht: (2026)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
von: You, Xiaoxing, et al.
Veröffentlicht: (2025)
Steering Away from Harm: An Adaptive Approach to Defending Vision Language Model Against Jailbreaks
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Quant Experts: Token-aware Adaptive Error Reconstruction with Mixture of Experts for Large Vision-Language Models Quantization
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
von: Jia, Chenwei, et al.
Veröffentlicht: (2026)
Adaptive Layer Selection for Efficient Vision Transformer Fine-Tuning
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
von: Devoto, Alessio, et al.
Veröffentlicht: (2024)
Dropout Prompt Learning: Towards Robust and Adaptive Vision-Language Models
von: Chen, Biao, et al.
Veröffentlicht: (2025)
von: Chen, Biao, et al.
Veröffentlicht: (2025)
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
Evolving Prompt Adaptation for Vision-Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
von: Zhang, Enming, et al.
Veröffentlicht: (2026)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
von: Zheng, Qi, et al.
Veröffentlicht: (2026)
Adaptive Residual-Update Steering for Low-Overhead Hallucination Mitigation in Large Vision Language Models
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
von: Zou, Zhengtao, et al.
Veröffentlicht: (2025)
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
von: Song, Fei, et al.
Veröffentlicht: (2025)
von: Song, Fei, et al.
Veröffentlicht: (2025)
Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings
von: Sharifzadeh, Sahand, et al.
Veröffentlicht: (2024)
von: Sharifzadeh, Sahand, et al.
Veröffentlicht: (2024)
KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
von: Xu, Wanshun, et al.
Veröffentlicht: (2025)
von: Xu, Wanshun, et al.
Veröffentlicht: (2025)
Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
von: He, Jialuo, et al.
Veröffentlicht: (2026)
von: He, Jialuo, et al.
Veröffentlicht: (2026)
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
von: Ryan, Yuriel, et al.
Veröffentlicht: (2026)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2026)
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
von: Im, Eun Woo, et al.
Veröffentlicht: (2025)
von: Im, Eun Woo, et al.
Veröffentlicht: (2025)
Is There Knowledge Left to Extract? Evidence of Fragility in Medically Fine-Tuned Vision-Language Models
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2026)
von: McLaughlin, Oliver, et al.
Veröffentlicht: (2026)
Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Ye, Wei, et al.
Veröffentlicht: (2024)
von: Ye, Wei, et al.
Veröffentlicht: (2024)
Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models
von: Korkmaz, Cansu, et al.
Veröffentlicht: (2025)
von: Korkmaz, Cansu, et al.
Veröffentlicht: (2025)
Adaptive Camera Sensor for Vision Models
von: Baek, Eunsu, et al.
Veröffentlicht: (2025)
von: Baek, Eunsu, et al.
Veröffentlicht: (2025)
TRIPS: Efficient Vision-and-Language Pre-training with Text-Relevant Image Patch Selection
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
von: Ma, Jie, et al.
Veröffentlicht: (2026)
von: Ma, Jie, et al.
Veröffentlicht: (2026)
STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language Models
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
von: Han, Yuhang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
von: Zhang, Xinsong, et al.
Veröffentlicht: (2025) -
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025) -
Vision Language Model-based Caption Evaluation Method Leveraging Visual Context Extraction
von: Maeda, Koki, et al.
Veröffentlicht: (2024) -
To Trust Or Not To Trust Your Vision-Language Model's Prediction
von: Dong, Hao, et al.
Veröffentlicht: (2025) -
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)