Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Zheng, Chen, Yiming, He, Nan, Chen, Jiahui, Li, Chaoyang, Qian, Houde, Sun, Lifeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SubFLOT: Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport
by: Jiang, Zheng, et al.
Published: (2026)
by: Jiang, Zheng, et al.
Published: (2026)
Decoupling Defense Strategies for Robust Image Watermarking
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
by: Li, Xuchen, et al.
Published: (2026)
by: Li, Xuchen, et al.
Published: (2026)
Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
by: Li, Shuang, et al.
Published: (2026)
by: Li, Shuang, et al.
Published: (2026)
Multi-Scale Global-Instance Prompt Tuning for Continual Test-time Adaptation in Medical Image Segmentation
by: Li, Lingrui, et al.
Published: (2026)
by: Li, Lingrui, et al.
Published: (2026)
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images
by: Li, Qingxiao, et al.
Published: (2026)
by: Li, Qingxiao, et al.
Published: (2026)
CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
by: Qian, Jiahe, et al.
Published: (2025)
by: Qian, Jiahe, et al.
Published: (2025)
GameIR: A Large-Scale Synthesized Ground-Truth Dataset for Image Restoration over Gaming Content
by: Zhou, Lebin, et al.
Published: (2024)
by: Zhou, Lebin, et al.
Published: (2024)
HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming
by: Chen, Jiahui, et al.
Published: (2026)
by: Chen, Jiahui, et al.
Published: (2026)
Guided Trajectory Optimization with Sparse Scaling for Test-Time Diffusion
by: Dai, Gang, et al.
Published: (2026)
by: Dai, Gang, et al.
Published: (2026)
SparseCoop: Cooperative Perception with Kinematic-Grounded Queries
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
CodePercept: Code-Grounded Visual STEM Perception for MLLMs
by: Guan, Tongkun, et al.
Published: (2026)
by: Guan, Tongkun, et al.
Published: (2026)
VOOM: Robust Visual Object Odometry and Mapping using Hierarchical Landmarks
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
GOReloc: Graph-based Object-Level Relocalization for Visual SLAM
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
A Survey on Open-Vocabulary Detection and Segmentation: Past, Present, and Future
by: Zhu, Chaoyang, et al.
Published: (2023)
by: Zhu, Chaoyang, et al.
Published: (2023)
Thyme: Think Beyond Images
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Zero-Shot Open-Vocabulary Human Motion Grounding with Test-Time Training
by: Zhou, Yunjiao, et al.
Published: (2025)
by: Zhou, Yunjiao, et al.
Published: (2025)
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence
by: Zhang, Yanbing, et al.
Published: (2026)
by: Zhang, Yanbing, et al.
Published: (2026)
Unified Representation Space for 3D Visual Grounding
by: Zheng, Yinuo, et al.
Published: (2025)
by: Zheng, Yinuo, et al.
Published: (2025)
TartanGround: A Large-Scale Dataset for Ground Robot Perception and Navigation
by: Patel, Manthan, et al.
Published: (2025)
by: Patel, Manthan, et al.
Published: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
by: Guo, Yuxiang, et al.
Published: (2025)
by: Guo, Yuxiang, et al.
Published: (2025)
One-Step Distillation of Discrete Diffusion Image Generators via Fixed-Point Iteration
by: Wang, Chaoyang, et al.
Published: (2026)
by: Wang, Chaoyang, et al.
Published: (2026)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
by: Avogaro, Niccolo, et al.
Published: (2026)
by: Avogaro, Niccolo, et al.
Published: (2026)
Start Small, Think Big: Curriculum-based Relative Policy Optimization for Visual Grounding
by: Yan, Qingyang, et al.
Published: (2025)
by: Yan, Qingyang, et al.
Published: (2025)
Visual Test-time Scaling for GUI Agent Grounding
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Super-Resolving Blurry Images with Events
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
NYC-Event-VPR: A Large-Scale High-Resolution Event-Based Visual Place Recognition Dataset in Dense Urban Environments
by: Pan, Taiyi, et al.
Published: (2024)
by: Pan, Taiyi, et al.
Published: (2024)
GGPT: Geometry Grounded Point Transformer
by: Chen, Yutong, et al.
Published: (2026)
by: Chen, Yutong, et al.
Published: (2026)
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
by: Qian, Yusu, et al.
Published: (2025)
by: Qian, Yusu, et al.
Published: (2025)
MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
by: Zheng, Lihao, et al.
Published: (2025)
by: Zheng, Lihao, et al.
Published: (2025)
ActFormer: Scalable Collaborative Perception via Active Queries
by: Huang, Suozhi, et al.
Published: (2024)
by: Huang, Suozhi, et al.
Published: (2024)
V-Thinker: Interactive Thinking with Images
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
Q-Ground: Image Quality Grounding with Large Multi-modality Models
by: Chen, Chaofeng, et al.
Published: (2024)
by: Chen, Chaofeng, et al.
Published: (2024)
Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
by: Chen, Zheng, et al.
Published: (2024)
by: Chen, Zheng, et al.
Published: (2024)
RESBev: Making BEV Perception More Robust
by: Zhuo, Lifeng, et al.
Published: (2026)
by: Zhuo, Lifeng, et al.
Published: (2026)
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation
by: He, Jingxuan, et al.
Published: (2026)
by: He, Jingxuan, et al.
Published: (2026)
Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
by: Ge, Jiusong, et al.
Published: (2026)
by: Ge, Jiusong, et al.
Published: (2026)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
Similar Items
-
SubFLOT: Submodel Extraction for Efficient and Personalized Federated Learning via Optimal Transport
by: Jiang, Zheng, et al.
Published: (2026) -
Decoupling Defense Strategies for Robust Image Watermarking
by: Chen, Jiahui, et al.
Published: (2026) -
Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
by: Li, Xuchen, et al.
Published: (2026) -
Disentangling to Re-couple: Resolving the Similarity-Controllability Paradox in Subject-Driven Text-to-Image Generation
by: Li, Shuang, et al.
Published: (2026) -
Multi-Scale Global-Instance Prompt Tuning for Continual Test-time Adaptation in Medical Image Segmentation
by: Li, Lingrui, et al.
Published: (2026)