NICE FACT: Diagnosing and Calibrating VLMs in Quantitative Reasoning for Kinematic Physics
Fuente:
arXiv
Salvato in:
| Autori principali: | Lan, Jian, Liu, Zhicheng, Wang, Xinpeng, Zhou, Yuhao, Chen, Haokun, Lv, Jiancheng, Plank, Barbara, Seidl, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
di: Zhou, Yuhao, et al.
Pubblicazione: (2026)
di: Zhou, Yuhao, et al.
Pubblicazione: (2026)
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
di: Lan, Jian, et al.
Pubblicazione: (2025)
di: Lan, Jian, et al.
Pubblicazione: (2025)
Style Quantization for Data-Efficient GAN Training
di: Wang, Jian, et al.
Pubblicazione: (2025)
di: Wang, Jian, et al.
Pubblicazione: (2025)
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
di: Lan, Jian, et al.
Pubblicazione: (2024)
di: Lan, Jian, et al.
Pubblicazione: (2024)
BINO: Encoder Centric Self Supervised Stereo With Native Pair Input
di: Zhou, Haokun
Pubblicazione: (2026)
di: Zhou, Haokun
Pubblicazione: (2026)
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
di: Lan, Jian, et al.
Pubblicazione: (2025)
di: Lan, Jian, et al.
Pubblicazione: (2025)
Target-Guided Bayesian Flow Networks for Quantitatively Constrained CAD Generation
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
di: Ni, Shuo, et al.
Pubblicazione: (2026)
di: Ni, Shuo, et al.
Pubblicazione: (2026)
GPS: Distilling Compact Memories via Grid-based Patch Sampling for Efficient Online Class-Incremental Learning
di: Ma, Mingchuan, et al.
Pubblicazione: (2025)
di: Ma, Mingchuan, et al.
Pubblicazione: (2025)
FACT: A Simple and Efficient Framework for Active Finetuning
di: Xu, Wenshuai, et al.
Pubblicazione: (2026)
di: Xu, Wenshuai, et al.
Pubblicazione: (2026)
Do VLMs Have Bad Eyes? Diagnosing Compositional Failures via Mechanistic Interpretability
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
di: Aravindan, Ashwath Vaithinathan, et al.
Pubblicazione: (2025)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
di: Liu, Xinpeng, et al.
Pubblicazione: (2023)
The Solution for the CVPR2024 NICE Image Captioning Challenge
di: Huang, Longfei, et al.
Pubblicazione: (2024)
di: Huang, Longfei, et al.
Pubblicazione: (2024)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
di: Liu, Peng, et al.
Pubblicazione: (2025)
di: Liu, Peng, et al.
Pubblicazione: (2025)
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
di: Hong, Hyesoo, et al.
Pubblicazione: (2026)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
di: Wang, Fanyi, et al.
Pubblicazione: (2025)
di: Wang, Fanyi, et al.
Pubblicazione: (2025)
Learning Locally, Revising Globally: Global Reviser for Federated Learning with Noisy Labels
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
di: Tian, Yuxin, et al.
Pubblicazione: (2024)
FACT: Feature Adaptive Continual-learning Tracker for Multiple Object Tracking
di: Song, Rongzihan, et al.
Pubblicazione: (2024)
di: Song, Rongzihan, et al.
Pubblicazione: (2024)
ARTeFACT: Benchmarking Segmentation Models on Diverse Analogue Media Damage
di: Ivanova, Daniela, et al.
Pubblicazione: (2024)
di: Ivanova, Daniela, et al.
Pubblicazione: (2024)
Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers
di: Wang, Jiancheng, et al.
Pubblicazione: (2026)
di: Wang, Jiancheng, et al.
Pubblicazione: (2026)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
di: Yu, Jiaao, et al.
Pubblicazione: (2025)
di: Yu, Jiaao, et al.
Pubblicazione: (2025)
DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs
di: Pan, Jiazhen, et al.
Pubblicazione: (2026)
di: Pan, Jiazhen, et al.
Pubblicazione: (2026)
DISSECT: Diagnosing Where Vision Ends and Language Priors Begin in Scientific VLMs
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
di: Kukreja, Dikshant, et al.
Pubblicazione: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
di: Li, Weiming, et al.
Pubblicazione: (2025)
di: Li, Weiming, et al.
Pubblicazione: (2025)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
di: Xie, Zhuofan, et al.
Pubblicazione: (2026)
di: Xie, Zhuofan, et al.
Pubblicazione: (2026)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Caption This, Reason That: VLMs Caught in the Middle
di: Weng, Zihan, et al.
Pubblicazione: (2025)
di: Weng, Zihan, et al.
Pubblicazione: (2025)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
di: Elmansoury, Sary, et al.
Pubblicazione: (2025)
di: Elmansoury, Sary, et al.
Pubblicazione: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
di: Huang, Irene, et al.
Pubblicazione: (2024)
di: Huang, Irene, et al.
Pubblicazione: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
di: Jin, Hongbo, et al.
Pubblicazione: (2026)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
di: Törtei, Brigitta Malagurski, et al.
Pubblicazione: (2025)
di: Törtei, Brigitta Malagurski, et al.
Pubblicazione: (2025)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
di: Malakouti, Sina, et al.
Pubblicazione: (2024)
di: Malakouti, Sina, et al.
Pubblicazione: (2024)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
di: Li, Yiwei, et al.
Pubblicazione: (2026)
di: Li, Yiwei, et al.
Pubblicazione: (2026)
DiffuSyn Bench: Evaluating Vision-Language Models on Real-World Complexities with Diffusion-Generated Synthetic Benchmarks
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
di: Zhou, Haokun, et al.
Pubblicazione: (2024)
VisualActBench: Can VLMs See and Act like a Human?
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
di: Alghisi, Simone, et al.
Pubblicazione: (2025)
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
di: Jian, Ai, et al.
Pubblicazione: (2025)
di: Jian, Ai, et al.
Pubblicazione: (2025)
PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models
di: Mak, Chak-Wing, et al.
Pubblicazione: (2026)
di: Mak, Chak-Wing, et al.
Pubblicazione: (2026)
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs
di: Li, Jiangyang, et al.
Pubblicazione: (2026)
di: Li, Jiangyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
di: Zhou, Yuhao, et al.
Pubblicazione: (2026) -
Unveiling the "Fairness Seesaw": Discovering and Mitigating Gender and Race Bias in Vision-Language Models
di: Lan, Jian, et al.
Pubblicazione: (2025) -
Style Quantization for Data-Efficient GAN Training
di: Wang, Jian, et al.
Pubblicazione: (2025) -
Mind the Uncertainty in Human Disagreement: Evaluating Discrepancies between Model Predictions and Human Responses in VQA
di: Lan, Jian, et al.
Pubblicazione: (2024) -
BINO: Encoder Centric Self Supervised Stereo With Native Pair Input
di: Zhou, Haokun
Pubblicazione: (2026)