Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Mingyuan, Li, Meitang, Yang, Jingcheng, Jiang, Jize, Yan, Kaizhuo, Li, Zhaoheng, Yu, Hanchao, Zhang, Minjia, Nahrstedt, Klara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
by: Chi, Banghao, et al.
Published: (2026)
by: Chi, Banghao, et al.
Published: (2026)
QoS-QoE Translation with Large Language Model
by: Yu, Yingjie, et al.
Published: (2026)
by: Yu, Yingjie, et al.
Published: (2026)
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
by: Yang, Jingcheng, et al.
Published: (2026)
by: Yang, Jingcheng, et al.
Published: (2026)
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
by: Pi, Xinyu, et al.
Published: (2024)
by: Pi, Xinyu, et al.
Published: (2024)
ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
by: Wang, Yuquan, et al.
Published: (2025)
by: Wang, Yuquan, et al.
Published: (2025)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
ACT360: An Efficient 360-Degree Action Detection and Summarization Framework for Mission-Critical Training and Debriefing
by: Tiwari, Aditi, et al.
Published: (2025)
by: Tiwari, Aditi, et al.
Published: (2025)
Performance Characterization of Containers in Edge Computing
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
How RL Unlocks the Aha Moment in Geometric Interleaved Reasoning
by: Zhang, Xiangxiang, et al.
Published: (2026)
by: Zhang, Xiangxiang, et al.
Published: (2026)
Understanding Aha Moments: from External Observations to Internal Mechanisms
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025)
by: Zheng, Haozhen, et al.
Published: (2025)
Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games
by: Zhang, Xiaoqing, et al.
Published: (2025)
by: Zhang, Xiaoqing, et al.
Published: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
by: Zhou, Kaiwen, et al.
Published: (2025)
by: Zhou, Kaiwen, et al.
Published: (2025)
Viewport-based Neural 360° Image Compression
by: Liao, Jingwei, et al.
Published: (2026)
by: Liao, Jingwei, et al.
Published: (2026)
TraceNet: Segment one thing efficiently
by: Wu, Mingyuan, et al.
Published: (2024)
by: Wu, Mingyuan, et al.
Published: (2024)
Report on the NSF Workshop on Sustainable Computing for Sustainability (NSF WSCS 2024)
by: Guérin, Roch, et al.
Published: (2024)
by: Guérin, Roch, et al.
Published: (2024)
Pseudo Dataset Generation for Out-of-Domain Multi-Camera View Recommendation
by: Lee, Kuan-Ying, et al.
Published: (2024)
by: Lee, Kuan-Ying, et al.
Published: (2024)
AquaScope: Reliable Underwater Image Transmission on Mobile Devices
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
by: Zhou, Hengguang, et al.
Published: (2025)
by: Zhou, Hengguang, et al.
Published: (2025)
Can Aha Moments Be Fake? Towards Quantifying Decorative and True Thinking in Chain-of-Thought
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
EcoLens: Leveraging Multi-Objective Bayesian Optimization for Energy-Efficient Video Processing on Edge Devices
by: Civjan, Benjamin, et al.
Published: (2025)
by: Civjan, Benjamin, et al.
Published: (2025)
Federated Transfer Learning with Task Personalization for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
Aha Moments and Continued Confusion: An Analysis of Threshold Concepts through Student Reflections in the ACRL Framework
by: Eva, Nicole C., et al.
Published: (2021)
by: Eva, Nicole C., et al.
Published: (2021)
From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
by: Ha, Rui, et al.
Published: (2025)
by: Ha, Rui, et al.
Published: (2025)
Adaptive Unknown Fault Detection and Few-Shot Continual Learning for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
Do VLMs Truly "Read" Candlesticks? A Multi-Scale Benchmark for Visual Stock Price Forecasting
by: Hu, Kaiqi, et al.
Published: (2026)
by: Hu, Kaiqi, et al.
Published: (2026)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
by: Gu, Yifeng, et al.
Published: (2025)
by: Gu, Yifeng, et al.
Published: (2025)
Unsichtbares wird zum Aha‐Erlebnis
by: Susanne Rehn‐Taube
Published: (2025)
by: Susanne Rehn‐Taube
Published: (2025)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
Functional Moments Regression
by: Li, Mingyuan, et al.
Published: (2026)
by: Li, Mingyuan, et al.
Published: (2026)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
by: Zhang, Jianshu, et al.
Published: (2026)
by: Zhang, Jianshu, et al.
Published: (2026)
Can Compressed LLMs Truly Act? An Empirical Evaluation of Agentic Capabilities in LLM Compression
by: Dong, Peijie, et al.
Published: (2025)
by: Dong, Peijie, et al.
Published: (2025)
Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification
by: Wan, Yuxuan, et al.
Published: (2026)
by: Wan, Yuxuan, et al.
Published: (2026)
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
by: Hu, Zhiyuan, et al.
Published: (2025)
by: Hu, Zhiyuan, et al.
Published: (2025)
Similar Items
-
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025) -
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
by: Wu, Mingyuan, et al.
Published: (2025) -
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
by: Chi, Banghao, et al.
Published: (2026) -
QoS-QoE Translation with Large Language Model
by: Yu, Yingjie, et al.
Published: (2026) -
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
by: Yang, Jingcheng, et al.
Published: (2026)