Take A Step Back: Rethinking the Two Stages in Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Mingyu, Cai, Jiting, Liu, Mingyu, Xu, Yue, Lu, Cewu, Li, Yong-Lu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023)
by: Wang, Ziyu, et al.
Published: (2023)
IPR-1: Interactive Physical Reasoner
by: Zhang, Mingyu, et al.
Published: (2025)
by: Zhang, Mingyu, et al.
Published: (2025)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026)
by: Kim, Keuntae, et al.
Published: (2026)
Low-Rank Similarity Mining for Multimodal Dataset Distillation
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
by: Xu, Xinyu, et al.
Published: (2024)
by: Xu, Xinyu, et al.
Published: (2024)
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
SAM-guided Pseudo Label Enhancement for Multi-modal 3D Semantic Segmentation
by: Yang, Mingyu, et al.
Published: (2025)
by: Yang, Mingyu, et al.
Published: (2025)
Homogeneous Dynamics Space for Heterogeneous Humans
by: Liu, Xinpeng, et al.
Published: (2024)
by: Liu, Xinpeng, et al.
Published: (2024)
DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
Kalib: Easy Hand-Eye Calibration with Reference Point Tracking
by: Tang, Tutian, et al.
Published: (2024)
by: Tang, Tutian, et al.
Published: (2024)
Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Trajectory-Consistent Calibration for Cache-Accelerated Diffusion Models
by: Liang, Mingyu, et al.
Published: (2026)
by: Liang, Mingyu, et al.
Published: (2026)
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
by: Zhao, Shijie, et al.
Published: (2025)
by: Zhao, Shijie, et al.
Published: (2025)
Revisit Human-Scene Interaction via Space Occupancy
by: Liu, Xinpeng, et al.
Published: (2023)
by: Liu, Xinpeng, et al.
Published: (2023)
Editable Noise Map Inversion: Encoding Target-image into Noise For High-Fidelity Image Manipulation
by: Kang, Mingyu, et al.
Published: (2025)
by: Kang, Mingyu, et al.
Published: (2025)
When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
PALUM: Part-based Attention Learning for Unified Motion Retargeting
by: Liu, Siqi, et al.
Published: (2026)
by: Liu, Siqi, et al.
Published: (2026)
ShapeBoost: Boosting Human Shape Estimation with Part-Based Parameterization and Clothing-Preserving Augmentation
by: Bian, Siyuan, et al.
Published: (2024)
by: Bian, Siyuan, et al.
Published: (2024)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
Revisiting Surgical Instrument Segmentation Without Human Intervention: A Graph Partitioning View
by: Sheng, Mingyu, et al.
Published: (2024)
by: Sheng, Mingyu, et al.
Published: (2024)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
LIDEA: Human-to-Robot Imitation Learning via Implicit Feature Distillation and Explicit Geometry Alignment
by: Xu, Yifu, et al.
Published: (2026)
by: Xu, Yifu, et al.
Published: (2026)
Semantic-Enriched Latent Visual Reasoning
by: Xu, Tianrun, et al.
Published: (2026)
by: Xu, Tianrun, et al.
Published: (2026)
Digital Gene: Learning about the Physical World through Analytic Concepts
by: Sun, Jianhua, et al.
Published: (2025)
by: Sun, Jianhua, et al.
Published: (2025)
E-comIQ-ZH: A Human-Aligned Dataset and Benchmark for Fine-Grained Evaluation of E-commerce Posters with Chain-of-Thought
by: Sun, Meiqi, et al.
Published: (2026)
by: Sun, Meiqi, et al.
Published: (2026)
DeepScan: A Training-Free Framework for Visually Grounded Reasoning in Large Vision-Language Models
by: Li, Yangfu, et al.
Published: (2026)
by: Li, Yangfu, et al.
Published: (2026)
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
by: Lu, Lidong, et al.
Published: (2025)
by: Lu, Lidong, et al.
Published: (2025)
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
by: Guo, Tengda, et al.
Published: (2026)
by: Guo, Tengda, et al.
Published: (2026)
ChartM$^3$: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
by: Xu, Duo, et al.
Published: (2025)
by: Xu, Duo, et al.
Published: (2025)
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
Take A Shortcut Back: Mitigating the Gradient Vanishing for Training Spiking Neural Networks
by: Guo, Yufei, et al.
Published: (2024)
by: Guo, Yufei, et al.
Published: (2024)
AMNCutter: Affinity-Attention-Guided Multi-View Normalized Cutter for Unsupervised Surgical Instrument Segmentation
by: Sheng, Mingyu, et al.
Published: (2024)
by: Sheng, Mingyu, et al.
Published: (2024)
Snakes and Ladders: Two Steps Up for VideoMamba
by: Lu, Hui, et al.
Published: (2024)
by: Lu, Hui, et al.
Published: (2024)
Stereo-Inertial Poser: Towards Metric-Accurate Shape-Aware Motion Capture Using Sparse IMUs and a Single Stereo Camera
by: Tang, Tutian, et al.
Published: (2026)
by: Tang, Tutian, et al.
Published: (2026)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
by: Chen, Yan, et al.
Published: (2025)
by: Chen, Yan, et al.
Published: (2025)
A Modern Take on Visual Relationship Reasoning for Grasp Planning
by: Rabino, Paolo, et al.
Published: (2024)
by: Rabino, Paolo, et al.
Published: (2024)
An Efficient Framework for Crediting Data Contributors of Diffusion Models
by: Lin, Chris, et al.
Published: (2024)
by: Lin, Chris, et al.
Published: (2024)
Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models
by: Xiong, Guangzhi, et al.
Published: (2026)
by: Xiong, Guangzhi, et al.
Published: (2026)
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
by: Lu, Jinda, et al.
Published: (2024)
by: Lu, Jinda, et al.
Published: (2024)
Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
by: Wei, Jiude, et al.
Published: (2025)
by: Wei, Jiude, et al.
Published: (2025)
Similar Items
-
Dancing with Still Images: Video Distillation via Static-Dynamic Disentanglement
by: Wang, Ziyu, et al.
Published: (2023) -
IPR-1: Interactive Physical Reasoner
by: Zhang, Mingyu, et al.
Published: (2025) -
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
by: Kim, Keuntae, et al.
Published: (2026) -
Low-Rank Similarity Mining for Multimodal Dataset Distillation
by: Xu, Yue, et al.
Published: (2024) -
DISCO: Embodied Navigation and Interaction via Differentiable Scene Semantics and Dual-level Control
by: Xu, Xinyu, et al.
Published: (2024)