Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Jiaying, Rao, Jinmeng, Chen, Kezhen, Guo, Xiaoyuan, Zhang, Yawen, Sun, Baochen, Yang, Carl, Yang, Jie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Higher Layers Need More LoRA Experts
von: Gao, Chongyang, et al.
Veröffentlicht: (2024)
von: Gao, Chongyang, et al.
Veröffentlicht: (2024)
IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues
von: Yang, Diji, et al.
Veröffentlicht: (2024)
von: Yang, Diji, et al.
Veröffentlicht: (2024)
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024)
von: Qi, Lu, et al.
Veröffentlicht: (2024)
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
von: Geigle, Gregor, et al.
Veröffentlicht: (2024)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Instruction-Following Evaluation of Large Vision-Language Models
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
von: Shiono, Daiki, et al.
Veröffentlicht: (2025)
PUMGPT: A Large Vision-Language Model for Product Understanding
von: Xue, Wei, et al.
Veröffentlicht: (2023)
von: Xue, Wei, et al.
Veröffentlicht: (2023)
SleepVLM: Explainable and Rule-Grounded Sleep Staging via a Vision-Language Model
von: Deng, Guifeng, et al.
Veröffentlicht: (2026)
von: Deng, Guifeng, et al.
Veröffentlicht: (2026)
Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
von: Lei, Xuanyu, et al.
Veröffentlicht: (2024)
von: Lei, Xuanyu, et al.
Veröffentlicht: (2024)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
von: Wang, Ye, et al.
Veröffentlicht: (2025)
von: Wang, Ye, et al.
Veröffentlicht: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
von: Matta, Shiho, et al.
Veröffentlicht: (2025)
von: Matta, Shiho, et al.
Veröffentlicht: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Fine-Grained Evaluation of Large Vision-Language Models in Autonomous Driving
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
von: Wu, Ruijia, et al.
Veröffentlicht: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
von: Le, Quang-Hung, et al.
Veröffentlicht: (2024)
Vision-Language Modeling in PET/CT for Visual Grounding of Positive Findings
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
von: Huemann, Zachary, et al.
Veröffentlicht: (2025)
VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
von: Luo, Lingxiao, et al.
Veröffentlicht: (2024)
von: Luo, Lingxiao, et al.
Veröffentlicht: (2024)
Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
von: Chen, Jiahe, et al.
Veröffentlicht: (2025)
Superpixel Semantics Representation and Pre-training for Vision-Language Task
von: Zhang, Siyu, et al.
Veröffentlicht: (2023)
von: Zhang, Siyu, et al.
Veröffentlicht: (2023)
ReVisionLLM: Recursive Vision-Language Model for Temporal Grounding in Hour-Long Videos
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
von: Hannan, Tanveer, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models as Evaluators in Path Planning
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
von: Aghzal, Mohamed, et al.
Veröffentlicht: (2024)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
von: Wu, Xuyang, et al.
Veröffentlicht: (2024)
Predicate Debiasing in Vision-Language Models Integration for Scene Graph Generation Enhancement
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Image-Based Geolocation Using Large Vision-Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
CrossGET: Cross-Guided Ensemble of Tokens for Accelerating Vision-Language Transformers
von: Shi, Dachuan, et al.
Veröffentlicht: (2023)
von: Shi, Dachuan, et al.
Veröffentlicht: (2023)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Higher Layers Need More LoRA Experts
von: Gao, Chongyang, et al.
Veröffentlicht: (2024) -
IM-RAG: Multi-Round Retrieval-Augmented Generation Through Learning Inner Monologues
von: Yang, Diji, et al.
Veröffentlicht: (2024) -
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024) -
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?
von: Geigle, Gregor, et al.
Veröffentlicht: (2024) -
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)