VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ghosal, Soumya Suvra, Kim, Youngeun, Li, Zhuowei, Chaudhry, Ritwick, Xu, Linghan, Zhang, Hongjing, Zablocki, Jakub, Xing, Yifan, Zhang, Qin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
by: Ghosal, Soumya Suvra, et al.
Published: (2024)
Enhancing Deep Neural Network Reliability with Refinement and Calibration
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
by: Hebbalaguppe, Ramya, et al.
Published: (2026)
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
by: Dogra, Atharvan, et al.
Published: (2025)
by: Dogra, Atharvan, et al.
Published: (2025)
LaRe: Latent Refocusing for Multimodal Reasoning
by: Ma, Jizheng, et al.
Published: (2025)
by: Ma, Jizheng, et al.
Published: (2025)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Echoes as Anchors: Probabilistic Costs and Attention Refocusing in LLM Reasoning
by: Hao, Zhuoyuan, et al.
Published: (2026)
by: Hao, Zhuoyuan, et al.
Published: (2026)
VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
by: Li, Lingxiao, et al.
Published: (2025)
by: Li, Lingxiao, et al.
Published: (2025)
Learning to Refocus with Video Diffusion Models
by: Tedla, SaiKiran, et al.
Published: (2025)
by: Tedla, SaiKiran, et al.
Published: (2025)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
by: Ghosal, Soumya Suvra, et al.
Published: (2025)
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
by: Dong, Qihua, et al.
Published: (2026)
by: Dong, Qihua, et al.
Published: (2026)
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
by: Wang, Ziyi, et al.
Published: (2025)
by: Wang, Ziyi, et al.
Published: (2025)
Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts
by: Liu, Xu, et al.
Published: (2026)
by: Liu, Xu, et al.
Published: (2026)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
by: Chehade, Mohamad, et al.
Published: (2025)
by: Chehade, Mohamad, et al.
Published: (2025)
Transfer Q Star: Principled Decoding for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2024)
by: Chakraborty, Souradip, et al.
Published: (2024)
Real-Time Visual Attribution Streaming in Thinking Model
by: Kang, Seil, et al.
Published: (2026)
by: Kang, Seil, et al.
Published: (2026)
AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
by: Li, Chengzu, et al.
Published: (2025)
by: Li, Chengzu, et al.
Published: (2025)
Mainstream Culture Refocused
by: Zhong, Xueping
Published: (2019)
by: Zhong, Xueping
Published: (2019)
Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
by: Shen, Ruolin, et al.
Published: (2025)
by: Shen, Ruolin, et al.
Published: (2025)
Open-World Dynamic Prompt and Continual Visual Representation Learning
by: Kim, Youngeun, et al.
Published: (2024)
by: Kim, Youngeun, et al.
Published: (2024)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
by: Xie, Yupeng, et al.
Published: (2025)
by: Xie, Yupeng, et al.
Published: (2025)
A New Cure Rate Model with Discrete and Multiple Exposures
by: Pal, Suvra
Published: (2024)
by: Pal, Suvra
Published: (2024)
StruVis: Enhancing Reasoning-based Text-to-Image Generation via Thinking with Structured Vision
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
by: Lyu, Yuanhuiyi, et al.
Published: (2026)
VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integration
by: Khaki, Saeed, et al.
Published: (2026)
by: Khaki, Saeed, et al.
Published: (2026)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
Threshold-Consistent Margin Loss for Open-World Deep Metric Learning
by: Zhang, Qin, et al.
Published: (2023)
by: Zhang, Qin, et al.
Published: (2023)
Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes
by: Wang, Yaoting, et al.
Published: (2024)
by: Wang, Yaoting, et al.
Published: (2024)
Robustifying Point Cloud Networks by Refocusing
by: Levi, Meir Yossef, et al.
Published: (2023)
by: Levi, Meir Yossef, et al.
Published: (2023)
DiffCamera: Arbitrary Refocusing on Images
by: Wang, Yiyang, et al.
Published: (2025)
by: Wang, Yiyang, et al.
Published: (2025)
Refocusing spacetimes need not be strongly refocusing
by: Bauermeister, Friedrich
Published: (2026)
by: Bauermeister, Friedrich
Published: (2026)
MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
by: Kim, Youngeun
Published: (2026)
by: Kim, Youngeun
Published: (2026)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
by: Ma, Jingkun, et al.
Published: (2024)
by: Ma, Jingkun, et al.
Published: (2024)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
VisTR: Visualizations as Representations for Time-series Table Reasoning
by: Hao, Jianing, et al.
Published: (2024)
by: Hao, Jianing, et al.
Published: (2024)
Tracer dynamics in an interacting active bath: fluctuations and energy partition
by: Sarkar, Ritwick, et al.
Published: (2025)
by: Sarkar, Ritwick, et al.
Published: (2025)
Emergent short-range repulsion for attractively coupled active particles
by: Sarkar, Ritwick, et al.
Published: (2024)
by: Sarkar, Ritwick, et al.
Published: (2024)
Similar Items
-
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2025) -
IntCoOp: Interpretability-Aware Vision-Language Prompt Tuning
by: Ghosal, Soumya Suvra, et al.
Published: (2024) -
PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks
by: Ghosal, Soumya Suvra, et al.
Published: (2024) -
Enhancing Deep Neural Network Reliability with Refinement and Calibration
by: Hebbalaguppe, Ramya, et al.
Published: (2026) -
Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away
by: Ghosal, Soumya Suvra, et al.
Published: (2026)