Slow Perception: Let's Perceive Geometric Figures Step-by-step
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Haoran, Yin, Youyang, Li, Yumeng, Wang, Jia, Zhao, Liang, Sun, Jianjian, Ge, Zheng, Zhang, Xiangyu, Jiang, Daxin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Perception in Reflection
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
Small Language Model Meets with Reinforced Vision Vocabulary
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
by: Chen, Jinyue, et al.
Published: (2024)
by: Chen, Jinyue, et al.
Published: (2024)
Focus Anywhere for Fine-grained Multi-page Document Understanding
by: Liu, Chenglong, et al.
Published: (2024)
by: Liu, Chenglong, et al.
Published: (2024)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
by: Han, Chunrui, et al.
Published: (2023)
by: Han, Chunrui, et al.
Published: (2023)
CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
by: Wei, Yana, et al.
Published: (2025)
by: Wei, Yana, et al.
Published: (2025)
Skinned Motion Retargeting with Dense Geometric Interaction Perception
by: Ye, Zijie, et al.
Published: (2024)
by: Ye, Zijie, et al.
Published: (2024)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
by: Xu, Guowei, et al.
Published: (2024)
by: Xu, Guowei, et al.
Published: (2024)
WebVR: Benchmarking Multimodal LLMs for WebPage Recreation from Videos via Human-Aligned Visual Rubrics
by: Dai, Yuhong, et al.
Published: (2026)
by: Dai, Yuhong, et al.
Published: (2026)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
by: Li, Dinging, et al.
Published: (2026)
by: Li, Dinging, et al.
Published: (2026)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
by: NextStep Team, et al.
Published: (2025)
by: NextStep Team, et al.
Published: (2025)
GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
by: Yan, Haolong, et al.
Published: (2025)
by: Yan, Haolong, et al.
Published: (2025)
Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
by: Liang, Zhanhao, et al.
Published: (2024)
by: Liang, Zhanhao, et al.
Published: (2024)
Step1X-Edit: A Practical Framework for General Image Editing
by: Liu, Shiyu, et al.
Published: (2025)
by: Liu, Shiyu, et al.
Published: (2025)
One-Step Diffusion for Perceptual Image Compression
by: Jia, Yiwen, et al.
Published: (2026)
by: Jia, Yiwen, et al.
Published: (2026)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
by: Xiang, Kun, et al.
Published: (2024)
by: Xiang, Kun, et al.
Published: (2024)
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs
by: Wang, Chongyu, et al.
Published: (2026)
by: Wang, Chongyu, et al.
Published: (2026)
LCV2I: Communication-Efficient and High-Performance Collaborative Perception Framework with Low-Resolution LiDAR
by: Feng, Xinxin, et al.
Published: (2025)
by: Feng, Xinxin, et al.
Published: (2025)
DreamLLM: Synergistic Multimodal Comprehension and Creation
by: Dong, Runpei, et al.
Published: (2023)
by: Dong, Runpei, et al.
Published: (2023)
Let's Reward Step-by-Step: Step-Aware Contrastive Alignment for Vision-Language Navigation in Continuous Environments
by: Li, Haoyuan, et al.
Published: (2026)
by: Li, Haoyuan, et al.
Published: (2026)
M-DocSum: Do LVLMs Genuinely Comprehend Interleaved Image-Text in Document Summarization?
by: Yan, Haolong, et al.
Published: (2025)
by: Yan, Haolong, et al.
Published: (2025)
Step1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D Assets
by: Li, Weiyu, et al.
Published: (2025)
by: Li, Weiyu, et al.
Published: (2025)
Uncertainty-Participation Context Consistency Learning for Semi-supervised Semantic Segmentation
by: Yin, Jianjian, et al.
Published: (2024)
by: Yin, Jianjian, et al.
Published: (2024)
Merlin:Empowering Multimodal LLMs with Foresight Minds
by: Yu, En, et al.
Published: (2023)
by: Yu, En, et al.
Published: (2023)
SlowPerception: Physical-World Latency Attack against Visual Perception in Autonomous Driving
by: Ma, Chen, et al.
Published: (2024)
by: Ma, Chen, et al.
Published: (2024)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
by: Wu, Yuhuan, et al.
Published: (2026)
by: Wu, Yuhuan, et al.
Published: (2026)
Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller
by: Zang, Chuanqi, et al.
Published: (2024)
by: Zang, Chuanqi, et al.
Published: (2024)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
by: Zhu, Xiaorong, et al.
Published: (2025)
by: Zhu, Xiaorong, et al.
Published: (2025)
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
by: Sun, Jiayin, et al.
Published: (2026)
by: Sun, Jiayin, et al.
Published: (2026)
Step-by-step Layered Design Generation
by: Khan, Faizan Farooq, et al.
Published: (2025)
by: Khan, Faizan Farooq, et al.
Published: (2025)
Self-Supervised Visual Preference Alignment
by: Zhu, Ke, et al.
Published: (2024)
by: Zhu, Ke, et al.
Published: (2024)
Highlight Every Step: Knowledge Distillation via Collaborative Teaching
by: Zhao, Haoran, et al.
Published: (2019)
by: Zhao, Haoran, et al.
Published: (2019)
One at a Time: Progressive Multi-step Volumetric Probability Learning for Reliable 3D Scene Perception
by: Li, Bohan, et al.
Published: (2023)
by: Li, Bohan, et al.
Published: (2023)
DFEN: Dual Feature Equalization Network for Medical Image Segmentation
by: Yin, Jianjian, et al.
Published: (2025)
by: Yin, Jianjian, et al.
Published: (2025)
Similar Items
-
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
by: Yu, En, et al.
Published: (2025) -
Perception in Reflection
by: Wei, Yana, et al.
Published: (2025) -
Small Language Model Meets with Reinforced Vision Vocabulary
by: Wei, Haoran, et al.
Published: (2024) -
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
by: Chen, Jinyue, et al.
Published: (2024) -
Focus Anywhere for Fine-grained Multi-page Document Understanding
by: Liu, Chenglong, et al.
Published: (2024)