Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Pan, Zhiyu, Wu, Yizheng, Hua, Jiashen, Feng, Junyi, Yan, Shaotian, Deng, Bing, Cao, Zhiguo, Ye, Jieping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
por: Liu, Kaiyuan, et al.
Publicado: (2025)
por: Liu, Kaiyuan, et al.
Publicado: (2025)
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
por: Yan, Shaotian, et al.
Publicado: (2025)
por: Yan, Shaotian, et al.
Publicado: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
por: Yan, Shaotian, et al.
Publicado: (2026)
por: Yan, Shaotian, et al.
Publicado: (2026)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
por: Liu, Junjie, et al.
Publicado: (2023)
por: Liu, Junjie, et al.
Publicado: (2023)
Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking
por: Jiang, Wei, et al.
Publicado: (2025)
por: Jiang, Wei, et al.
Publicado: (2025)
Self-Supervised Class-Agnostic Motion Prediction with Spatial and Temporal Consistency Regularizations
por: Wang, Kewei, et al.
Publicado: (2024)
por: Wang, Kewei, et al.
Publicado: (2024)
Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning
por: Huang, Chenxi, et al.
Publicado: (2025)
por: Huang, Chenxi, et al.
Publicado: (2025)
On the Step Length Confounding in LLM Reasoning Data Selection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
SalaMAnder: Shapley-based Mathematical Expression Attribution and Metric for Chain-of-Thought Reasoning
por: Xin, Yue, et al.
Publicado: (2025)
por: Xin, Yue, et al.
Publicado: (2025)
Pseudo-Labeling by Multi-Policy Viewfinder Network for Image Cropping
por: Pan, Zhiyu, et al.
Publicado: (2024)
por: Pan, Zhiyu, et al.
Publicado: (2024)
Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
por: Sun, Shengyang, et al.
Publicado: (2026)
por: Sun, Shengyang, et al.
Publicado: (2026)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
por: Ge, Yuyao, et al.
Publicado: (2025)
por: Ge, Yuyao, et al.
Publicado: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
por: Shen, Yifan, et al.
Publicado: (2025)
por: Shen, Yifan, et al.
Publicado: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
por: Berman, Shmuel, et al.
Publicado: (2025)
por: Berman, Shmuel, et al.
Publicado: (2025)
Instance Consistency Regularization for Semi-Supervised 3D Instance Segmentation
por: Wu, Yizheng, et al.
Publicado: (2024)
por: Wu, Yizheng, et al.
Publicado: (2024)
Enhancing Spatial Reasoning through Visual and Textual Thinking
por: Liang, Xun, et al.
Publicado: (2025)
por: Liang, Xun, et al.
Publicado: (2025)
Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
por: Liu, Kaiyuan, et al.
Publicado: (2025)
por: Liu, Kaiyuan, et al.
Publicado: (2025)
Instance-adaptive Zero-shot Chain-of-Thought Prompting
por: Yuan, Xiaosong, et al.
Publicado: (2024)
por: Yuan, Xiaosong, et al.
Publicado: (2024)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
por: Ye, Junyi, et al.
Publicado: (2024)
por: Ye, Junyi, et al.
Publicado: (2024)
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
por: Zhang, Xiaofeng, et al.
Publicado: (2024)
AnyID: Ultra-Fidelity Universal Identity-Preserving Video Generation from Any Visual References
por: Wang, Jiahao, et al.
Publicado: (2026)
por: Wang, Jiahao, et al.
Publicado: (2026)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
por: Gambashidze, Alexander, et al.
Publicado: (2025)
por: Gambashidze, Alexander, et al.
Publicado: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
por: Zhang, Yuyou, et al.
Publicado: (2025)
por: Zhang, Yuyou, et al.
Publicado: (2025)
Institutional Trust and the Domestic AI Advantage: Evidence from DeepSeek and ChatGPT Users in China
por: Huang, Jiashen, et al.
Publicado: (2026)
por: Huang, Jiashen, et al.
Publicado: (2026)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
por: Mayer, Julius, et al.
Publicado: (2025)
por: Mayer, Julius, et al.
Publicado: (2025)
CT3D++: Improving 3D Object Detection with Keypoint-induced Channel-wise Transformer
por: Sheng, Hualian, et al.
Publicado: (2024)
por: Sheng, Hualian, et al.
Publicado: (2024)
EchoShot: Multi-Shot Portrait Video Generation
por: Wang, Jiahao, et al.
Publicado: (2025)
por: Wang, Jiahao, et al.
Publicado: (2025)
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute
por: Liu, Sheng, et al.
Publicado: (2025)
por: Liu, Sheng, et al.
Publicado: (2025)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
por: Cui, Jiahao, et al.
Publicado: (2024)
por: Cui, Jiahao, et al.
Publicado: (2024)
Geometry-aware Reconstruction and Fusion-refined Rendering for Generalizable Neural Radiance Fields
por: Liu, Tianqi, et al.
Publicado: (2024)
por: Liu, Tianqi, et al.
Publicado: (2024)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
por: Elmansoury, Sary, et al.
Publicado: (2025)
por: Elmansoury, Sary, et al.
Publicado: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
por: Törtei, Brigitta Malagurski, et al.
Publicado: (2025)
por: Törtei, Brigitta Malagurski, et al.
Publicado: (2025)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
por: Khezresmaeilzadeh, Tina, et al.
Publicado: (2026)
por: Khezresmaeilzadeh, Tina, et al.
Publicado: (2026)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
por: Chen, Kaitao, et al.
Publicado: (2025)
por: Chen, Kaitao, et al.
Publicado: (2025)
Probing Visual Language Priors in VLMs
por: Luo, Tiange, et al.
Publicado: (2024)
por: Luo, Tiange, et al.
Publicado: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
por: Zhu, He, et al.
Publicado: (2025)
por: Zhu, He, et al.
Publicado: (2025)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
por: Li, Qiaoru, et al.
Publicado: (2026)
por: Li, Qiaoru, et al.
Publicado: (2026)
Ejemplares similares
-
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
por: Liu, Kaiyuan, et al.
Publicado: (2025) -
Don't Take Things Out of Context: Attention Intervention for Enhancing Chain-of-Thought Reasoning in Large Language Models
por: Yan, Shaotian, et al.
Publicado: (2025) -
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
por: Yan, Shaotian, et al.
Publicado: (2026) -
Concise and Organized Perception Facilitates Reasoning in Large Language Models
por: Liu, Junjie, et al.
Publicado: (2023) -
Semi-Supervised High Dynamic Range Image Reconstructing via Bi-Level Uncertain Area Masking
por: Jiang, Wei, et al.
Publicado: (2025)