PathGLS: Evaluating Pathology Vision-Language Models without Ground Truth through Multi-Dimensional Consistency
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Minbing, Meng, Zhu, Su, Fei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pathological Truth Bias in Vision-Language Models
von: Thube, Yash
Veröffentlicht: (2025)
von: Thube, Yash
Veröffentlicht: (2025)
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
von: Duan, Jinhao, et al.
Veröffentlicht: (2025)
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026)
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
von: Yuan, Ruicheng, et al.
Veröffentlicht: (2026)
von: Yuan, Ruicheng, et al.
Veröffentlicht: (2026)
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
von: Gao, Jensen, et al.
Veröffentlicht: (2023)
Self-Supervised Multi-Object Tracking with Path Consistency
von: Lu, Zijia, et al.
Veröffentlicht: (2024)
von: Lu, Zijia, et al.
Veröffentlicht: (2024)
GLS: Geometry-aware 3D Language Gaussian Splatting
von: Qiu, Jiaxiong, et al.
Veröffentlicht: (2024)
von: Qiu, Jiaxiong, et al.
Veröffentlicht: (2024)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
von: Chen, Boyuan, et al.
Veröffentlicht: (2026)
von: Chen, Boyuan, et al.
Veröffentlicht: (2026)
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
von: Xue, Haotian, et al.
Veröffentlicht: (2025)
Towards Self-Refinement of Vision-Language Models with Triangular Consistency
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
von: Deng, Yunlong, et al.
Veröffentlicht: (2025)
MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
Detecting Performance Degradation under Data Shift in Pathology Vision-Language Model
von: Guan, Hao, et al.
Veröffentlicht: (2026)
von: Guan, Hao, et al.
Veröffentlicht: (2026)
Efficient and Comprehensive Feature Extraction in Large Vision-Language Model for Pathology Analysis
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
von: Zhang, Shengxuming, et al.
Veröffentlicht: (2024)
VTCBench: Can Vision-Language Models Understand Long Context with Vision-Text Compression?
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2025)
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
von: Qu, Xiaoye, et al.
Veröffentlicht: (2024)
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
von: Back, Kyungryul, et al.
Veröffentlicht: (2025)
First Multi-Dimensional Evaluation of Flowchart Comprehension for Multimodal Large Language Models
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
von: Zhang, Enming, et al.
Veröffentlicht: (2024)
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework
von: Han, Xiao, et al.
Veröffentlicht: (2024)
von: Han, Xiao, et al.
Veröffentlicht: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
von: Wan, David, et al.
Veröffentlicht: (2024)
von: Wan, David, et al.
Veröffentlicht: (2024)
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
von: Chen, Kaitao, et al.
Veröffentlicht: (2024)
von: Chen, Kaitao, et al.
Veröffentlicht: (2024)
MM-MoralBench: A MultiModal Moral Evaluation Benchmark for Large Vision-Language Models
von: Yan, Bei, et al.
Veröffentlicht: (2024)
von: Yan, Bei, et al.
Veröffentlicht: (2024)
Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
TruthLens: Visual Grounding for Universal DeepFake Reasoning
von: Kundu, Rohit, et al.
Veröffentlicht: (2025)
von: Kundu, Rohit, et al.
Veröffentlicht: (2025)
HMGIE: Hierarchical and Multi-Grained Inconsistency Evaluation for Vision-Language Data Cleansing
von: Zhu, Zihao, et al.
Veröffentlicht: (2024)
von: Zhu, Zihao, et al.
Veröffentlicht: (2024)
Self-Evolving Spatial Reasoning in Vision Language Models via Geometric Logic Consistency
von: Liu, Junming, et al.
Veröffentlicht: (2026)
von: Liu, Junming, et al.
Veröffentlicht: (2026)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhao, et al.
Veröffentlicht: (2026)
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
von: Yang, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zhiyuan, et al.
Veröffentlicht: (2026)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
von: Li, Rang, et al.
Veröffentlicht: (2025)
von: Li, Rang, et al.
Veröffentlicht: (2025)
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
von: Hua, Shengyi, et al.
Veröffentlicht: (2025)
von: Hua, Shengyi, et al.
Veröffentlicht: (2025)
Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
von: Kumbhar, Shrinidhi, et al.
Veröffentlicht: (2026)
Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
von: Ernhofer, Benjamin Raphael, et al.
Veröffentlicht: (2025)
To Agree or To Be Right? The Grounding-Sycophancy Tradeoff in Medical Vision-Language Models
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
von: Aranya, OFM Riaz Rahman, et al.
Veröffentlicht: (2026)
Echo-Path: Pathology-Conditioned Echo Video Generation
von: Muhammad, Kabir Hamzah, et al.
Veröffentlicht: (2025)
von: Muhammad, Kabir Hamzah, et al.
Veröffentlicht: (2025)
Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
von: Yu, Keunwoo Peter, et al.
Veröffentlicht: (2025)
Towards Efficient and General-Purpose Few-Shot Misclassification Detection for Vision-Language Models
von: Zeng, Fanhu, et al.
Veröffentlicht: (2025)
von: Zeng, Fanhu, et al.
Veröffentlicht: (2025)
Harnessing Large Vision and Language Models in Agriculture: A Review
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
von: Zhu, Hongyan, et al.
Veröffentlicht: (2024)
Multi-Modal Foundation Models for Computational Pathology: A Survey
von: Li, Dong, et al.
Veröffentlicht: (2025)
von: Li, Dong, et al.
Veröffentlicht: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
von: Cheng, Zihui, et al.
Veröffentlicht: (2024)
TinyLVLM-eHub: Towards Comprehensive and Efficient Evaluation for Large Vision-Language Models
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
von: Shao, Wenqi, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Pathological Truth Bias in Vision-Language Models
von: Thube, Yash
Veröffentlicht: (2025) -
TruthPrInt: Mitigating Large Vision-Language Models Object Hallucination Via Latent Truthful-Guided Pre-Intervention
von: Duan, Jinhao, et al.
Veröffentlicht: (2025) -
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
von: Dong, Xinpeng, et al.
Veröffentlicht: (2026) -
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
von: Yuan, Ruicheng, et al.
Veröffentlicht: (2026) -
Physically Grounded Vision-Language Models for Robotic Manipulation
von: Gao, Jensen, et al.
Veröffentlicht: (2023)