Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Yifan, Liu, Yuanzhe, Zhu, Jingyuan, Cao, Xu, Zhang, Xiaofeng, He, Yixiao, Ye, Wenming, Rehg, James Matthew, Lourentzou, Ismini |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)
by: Holla, Meghana, et al.
Published: (2023)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
by: Shen, Ying, et al.
Published: (2023)
by: Shen, Ying, et al.
Published: (2023)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
by: Venkatesh, Kavana, et al.
Published: (2024)
by: Venkatesh, Kavana, et al.
Published: (2024)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026)
by: Ogunleye, Makanjuola, et al.
Published: (2026)
Iterative Reasoning Preference Optimization
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
by: Pang, Richard Yuanzhe, et al.
Published: (2024)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025)
by: Li, Xinzhuo, et al.
Published: (2025)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026)
by: Yu, Tianjiao, et al.
Published: (2026)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
by: Wang, Yeyuan, et al.
Published: (2025)
by: Wang, Yeyuan, et al.
Published: (2025)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
by: Pan, Zhiyu, et al.
Published: (2026)
by: Pan, Zhiyu, et al.
Published: (2026)
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
by: Liu, Yuanzhe, et al.
Published: (2026)
by: Liu, Yuanzhe, et al.
Published: (2026)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
by: He, Jixuan, et al.
Published: (2026)
by: He, Jixuan, et al.
Published: (2026)
Fine-Grained Self-Endorsement Improves Factuality and Reasoning
by: Wang, Ante, et al.
Published: (2024)
by: Wang, Ante, et al.
Published: (2024)
Why Is Spatial Reasoning Hard for VLMs? An Attention Mechanism Perspective on Focus Areas
by: Chen, Shiqi, et al.
Published: (2025)
by: Chen, Shiqi, et al.
Published: (2025)
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
by: Nguyen, Kiet A., et al.
Published: (2024)
by: Nguyen, Kiet A., et al.
Published: (2024)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
by: Zhu, Wenhong, et al.
Published: (2024)
by: Zhu, Wenhong, et al.
Published: (2024)
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs
by: Zhu, Mengdan, et al.
Published: (2026)
by: Zhu, Mengdan, et al.
Published: (2026)
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
by: Chaubey, Ashutosh, et al.
Published: (2026)
by: Chaubey, Ashutosh, et al.
Published: (2026)
Spatial Hierarchy and Temporal Attention Guided Cross Masking for Self-supervised Skeleton-based Action Recognition
by: Yin, Xinpeng, et al.
Published: (2024)
by: Yin, Xinpeng, et al.
Published: (2024)
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
by: Sayem, MD Khalequzzaman Chowdhury, et al.
Published: (2026)
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
by: Fan, Sinan, et al.
Published: (2025)
by: Fan, Sinan, et al.
Published: (2025)
Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
by: Sun, Zhen, et al.
Published: (2025)
by: Sun, Zhen, et al.
Published: (2025)
YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning
by: Le, Yifan
Published: (2026)
by: Le, Yifan
Published: (2026)
LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
by: Shen, Ying, et al.
Published: (2025)
by: Shen, Ying, et al.
Published: (2025)
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
by: Pan, Zhenyu, et al.
Published: (2025)
by: Pan, Zhenyu, et al.
Published: (2025)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024)
by: Wahed, Muntasir, et al.
Published: (2024)
Coarse-Grained Dynamics with Spatial Disorder and Non-Markovian Memory
by: Liu, Chuyi, et al.
Published: (2026)
by: Liu, Chuyi, et al.
Published: (2026)
Fine-grained Video Dubbing Duration Alignment with Segment Supervised Preference Optimization
by: Cui, Chaoqun, et al.
Published: (2025)
by: Cui, Chaoqun, et al.
Published: (2025)
KV-Compress: Paged KV-Cache Compression with Variable Compression Rates per Attention Head
by: Rehg, Isaac
Published: (2024)
by: Rehg, Isaac
Published: (2024)
Improving Fine-grained Visual Understanding in VLMs through Text-Only Training
by: Choi, Dasol, et al.
Published: (2024)
by: Choi, Dasol, et al.
Published: (2024)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
by: Wu, Ziyi, et al.
Published: (2025)
by: Wu, Ziyi, et al.
Published: (2025)
Arena as Offline Reward: Efficient Fine-Grained Preference Optimization for Diffusion Models
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
Divide, Optimize, Merge: Fine-Grained LLM Agent Optimization at Scale
by: Liu, Jiale, et al.
Published: (2025)
by: Liu, Jiale, et al.
Published: (2025)
Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
Similar Items
-
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025) -
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023) -
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026) -
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026) -
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
by: Shen, Ying, et al.
Published: (2023)