CARPE: Context-Aware Image Representation Prioritization via Ensemble for Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Donghee, Cai, Rui, Zhao, Zhe |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
von: Zhou, Zijie, et al.
Veröffentlicht: (2026)
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024)
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
von: Wang, Xinhao, et al.
Veröffentlicht: (2026)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
von: Song, Fei, et al.
Veröffentlicht: (2025)
von: Song, Fei, et al.
Veröffentlicht: (2025)
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
von: Cai, Rui, et al.
Veröffentlicht: (2025)
von: Cai, Rui, et al.
Veröffentlicht: (2025)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
From Captions to Rewards (CAREVL): Leveraging Large Language Model Experts for Enhanced Reward Modeling in Large Vision-Language Models
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
von: Dai, Muzhi, et al.
Veröffentlicht: (2025)
Image Tokens Matter: Mitigating Hallucination in Discrete Tokenizer-based Large Vision-Language Models via Latent Editing
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
von: Wang, Weixing, et al.
Veröffentlicht: (2025)
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
von: Wang, Yuheng, et al.
Veröffentlicht: (2026)
In-Context Learning Improves Compositional Understanding of Vision-Language Models
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
von: Nulli, Matteo, et al.
Veröffentlicht: (2024)
LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model
von: Xue, Yuan, et al.
Veröffentlicht: (2024)
von: Xue, Yuan, et al.
Veröffentlicht: (2024)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Attention Prompting on Image for Large Vision-Language Models
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
von: Yu, Runpeng, et al.
Veröffentlicht: (2024)
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
von: Datta, Shounak, et al.
Veröffentlicht: (2025)
von: Datta, Shounak, et al.
Veröffentlicht: (2025)
Fine-Grained Post-Training Quantization for Large Vision Language Models with Quantization-Aware Integrated Gradients
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
von: Xiang, Ziwei, et al.
Veröffentlicht: (2026)
Residual Decoding: Mitigating Hallucinations in Large Vision-Language Models via History-Aware Residual Guidance
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
von: Chen, Xinrong, et al.
Veröffentlicht: (2026)
Perceptio: Perception Enhanced Vision Language Models via Spatial Token Generation
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
von: Li, Yuchen, et al.
Veröffentlicht: (2026)
Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
von: Panchal, Utsav, et al.
Veröffentlicht: (2025)
von: Panchal, Utsav, et al.
Veröffentlicht: (2025)
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
von: Zhang, Juntian, et al.
Veröffentlicht: (2025)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
von: Shalabi, Fatma, et al.
Veröffentlicht: (2024)
Prescribing the Right Remedy: Mitigating Hallucinations in Large Vision-Language Models via Targeted Instruction Tuning
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
von: Min, Kyungmin, et al.
Veröffentlicht: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
Medical Large Vision Language Models with Multi-Image Visual Ability
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
The Geometry of Representational Failures in Vision Language Models
von: Savietto, Daniele, et al.
Veröffentlicht: (2026)
von: Savietto, Daniele, et al.
Veröffentlicht: (2026)
PATIMT-Bench: A Multi-Scenario Benchmark for Position-Aware Text Image Machine Translation in Large Vision-Language Models
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
von: Zhuang, Wanru, et al.
Veröffentlicht: (2025)
Cross-Cultural Value Awareness in Large Vision-Language Models
von: Howard, Phillip, et al.
Veröffentlicht: (2026)
von: Howard, Phillip, et al.
Veröffentlicht: (2026)
GA-VLN: Geometry-Aware BEV Representation for Efficient Vision-Language Navigation
von: Yang, Jiahao, et al.
Veröffentlicht: (2026)
von: Yang, Jiahao, et al.
Veröffentlicht: (2026)
MedVH: Towards Systematic Evaluation of Hallucination for Large Vision Language Models in the Medical Context
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
Compound Expression Recognition via Large Vision-Language Models
von: Yu, Jun, et al.
Veröffentlicht: (2025)
von: Yu, Jun, et al.
Veröffentlicht: (2025)
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
von: Yang, Siyuan, et al.
Veröffentlicht: (2026)
von: Yang, Siyuan, et al.
Veröffentlicht: (2026)
Phantasia: Context-Adaptive Backdoors in Vision Language Models
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
von: Tran, Nam Duong, et al.
Veröffentlicht: (2026)
Hidden Clones: Exposing and Fixing Family Bias in Vision-Language Model Ensembles
von: Bugaud, Zacharie
Veröffentlicht: (2026)
von: Bugaud, Zacharie
Veröffentlicht: (2026)
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
von: Zhang, Zilun, et al.
Veröffentlicht: (2025)
von: Zhang, Zilun, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Evaluation for Vision-Language Models
von: Kostumov, Vasily, et al.
Veröffentlicht: (2024)
von: Kostumov, Vasily, et al.
Veröffentlicht: (2024)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2024)
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models
von: Zhou, Zijie, et al.
Veröffentlicht: (2026) -
Large Vision-Language Models as Emotion Recognizers in Context Awareness
von: Lei, Yuxuan, et al.
Veröffentlicht: (2024) -
QAPruner: Quantization-Aware Vision Token Pruning for Multimodal Large Language Models
von: Wang, Xinhao, et al.
Veröffentlicht: (2026) -
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
von: Na, Youngjin, et al.
Veröffentlicht: (2025) -
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
von: Song, Fei, et al.
Veröffentlicht: (2025)