GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
Fuente:
arXiv
Saved in:
| Main Authors: | Zhan, Yufei, Wu, Ziheng, Zhu, Yousong, Xue, Rongkun, Luo, Ruipu, Chen, Zhenghao, Zhang, Can, Li, Yifan, He, Zhentao, Yang, Zheming, Tang, Ming, Qiu, Minghui, Wang, Jinqiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025)
by: Yu, Jiachen, et al.
Published: (2025)
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
by: Wu, Ziheng, et al.
Published: (2025)
by: Wu, Ziheng, et al.
Published: (2025)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025)
by: He, Zhentao, et al.
Published: (2025)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
by: Zheng, Shurong, et al.
Published: (2026)
by: Zheng, Shurong, et al.
Published: (2026)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
by: Zhan, Yufei, et al.
Published: (2025)
by: Zhan, Yufei, et al.
Published: (2025)
Griffon: Spelling out All Object Locations at Any Granularity with Large Language Models
by: Zhan, Yufei, et al.
Published: (2023)
by: Zhan, Yufei, et al.
Published: (2023)
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
by: Li, Qingbin, et al.
Published: (2025)
by: Li, Qingbin, et al.
Published: (2025)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
by: Li, Zejun, et al.
Published: (2024)
by: Li, Zejun, et al.
Published: (2024)
FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
by: Yang, Fan, et al.
Published: (2026)
by: Yang, Fan, et al.
Published: (2026)
Valley: Video Assistant with Large Language model Enhanced abilitY
by: Luo, Ruipu, et al.
Published: (2023)
by: Luo, Ruipu, et al.
Published: (2023)
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
by: Zhou, Guanghao, et al.
Published: (2025)
by: Zhou, Guanghao, et al.
Published: (2025)
Ten computational challenges in human virome studies
by: Wu, Yifan, et al.
Published: (2024)
by: Wu, Yifan, et al.
Published: (2024)
From Seeing to Predicting: A Vision-Language Framework for Trajectory Forecasting and Controlled Video Generation
by: Yang, Fan, et al.
Published: (2025)
by: Yang, Fan, et al.
Published: (2025)
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
by: Wang, Yongqi, et al.
Published: (2026)
by: Wang, Yongqi, et al.
Published: (2026)
Efficient Masked Autoencoders with Self-Consistency
by: Li, Zhaowen, et al.
Published: (2023)
by: Li, Zhaowen, et al.
Published: (2023)
ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control
by: Tang, Zhentao, et al.
Published: (2026)
by: Tang, Zhentao, et al.
Published: (2026)
Neural Conformal Control for Time Series Forecasting
by: Li, Ruipu, et al.
Published: (2024)
by: Li, Ruipu, et al.
Published: (2024)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection
by: Qian, Long, et al.
Published: (2025)
by: Qian, Long, et al.
Published: (2025)
Mixture-of-Retrieval Experts for Reasoning-Guided Multimodal Knowledge Exploitation
by: Peng, Chunyi, et al.
Published: (2025)
by: Peng, Chunyi, et al.
Published: (2025)
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
by: Zhang, Yiming, et al.
Published: (2026)
by: Zhang, Yiming, et al.
Published: (2026)
AnyPattern: Towards In-context Image Copy Detection
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
by: Liang, Zichen, et al.
Published: (2025)
by: Liang, Zichen, et al.
Published: (2025)
SEEKR: Selective Attention-Guided Knowledge Retention for Continual Learning of Large Language Models
by: He, Jinghan, et al.
Published: (2024)
by: He, Jinghan, et al.
Published: (2024)
LightPlanner: Unleashing the Reasoning Capabilities of Lightweight Large Language Models in Task Planning
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
What's in the Box? Reasoning about Unseen Objects from Multimodal Cues
by: Ying, Lance, et al.
Published: (2025)
by: Ying, Lance, et al.
Published: (2025)
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
by: Tian, Yuhao, et al.
Published: (2025)
by: Tian, Yuhao, et al.
Published: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
by: Zhou, Qianrui, et al.
Published: (2025)
by: Zhou, Qianrui, et al.
Published: (2025)
FLARE: A Framework for Stellar Flare Forecasting using Stellar Physical Properties and Historical Records
by: Zhu, Bingke, et al.
Published: (2025)
by: Zhu, Bingke, et al.
Published: (2025)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
by: An, Hongyan, et al.
Published: (2025)
by: An, Hongyan, et al.
Published: (2025)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
by: Qi, Yukun, et al.
Published: (2026)
by: Qi, Yukun, et al.
Published: (2026)
A high-order accurate moving mesh finite element method for the radial Kohn--Sham equation
by: Luo, Zheming, et al.
Published: (2024)
by: Luo, Zheming, et al.
Published: (2024)
Invariant theory of $\imath$quantum groups of type AIII
by: Luo, Li, et al.
Published: (2023)
by: Luo, Li, et al.
Published: (2023)
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition
by: Cong, Kaixuan, et al.
Published: (2025)
by: Cong, Kaixuan, et al.
Published: (2025)
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
by: Luo, Sha, et al.
Published: (2026)
by: Luo, Sha, et al.
Published: (2026)
Similar Items
-
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
by: Li, Yifan, et al.
Published: (2025) -
VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories?
by: Yu, Jiachen, et al.
Published: (2025) -
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design
by: Wu, Ziheng, et al.
Published: (2025) -
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
by: He, Zhentao, et al.
Published: (2025) -
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2025)