Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Jialong, Zhang, Xiaoying, Yuan, Hongyi, Zhang, Xiangcheng, Huang, Tianhao, He, Changjing, Deng, Chaoyi, Zhang, Renrui, Wu, Youbin, Long, Mingsheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025)
by: Zhang, Xinchen, et al.
Published: (2025)
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024)
by: Deng, Chaoyi, et al.
Published: (2024)
RLVR-World: Training World Models with Reinforcement Learning
by: Wu, Jialong, et al.
Published: (2025)
by: Wu, Jialong, et al.
Published: (2025)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025)
by: Huang, Siqiao, et al.
Published: (2025)
Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
by: Zhong, Jincheng, et al.
Published: (2025)
by: Zhong, Jincheng, et al.
Published: (2025)
Foresight Diffusion: Improving Sampling Consistency in Predictive Diffusion Models
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Trajectory World Models for Heterogeneous Environments
by: Yin, Shaofeng, et al.
Published: (2025)
by: Yin, Shaofeng, et al.
Published: (2025)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Adaptive Ability Decomposing for Unlocking Large Reasoning Model Effective Reinforcement Learning
by: Chen, Zhipeng, et al.
Published: (2026)
by: Chen, Zhipeng, et al.
Published: (2026)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
by: Huang, Yuesheng, et al.
Published: (2025)
by: Huang, Yuesheng, et al.
Published: (2025)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
HarmonyDream: Task Harmonization Inside World Models
by: Ma, Haoyu, et al.
Published: (2023)
by: Ma, Haoyu, et al.
Published: (2023)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Learning to Adapt SFT Data for Better Reasoning Generalization
by: Sun, Lisong, et al.
Published: (2026)
by: Sun, Lisong, et al.
Published: (2026)
A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets
by: Jia, Zexi, et al.
Published: (2025)
by: Jia, Zexi, et al.
Published: (2025)
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
by: Chen, Mingrui, et al.
Published: (2025)
by: Chen, Mingrui, et al.
Published: (2025)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Assessing the Human Likeness of AI-Generated Counterspeech
by: Song, Xiaoying, et al.
Published: (2024)
by: Song, Xiaoying, et al.
Published: (2024)
Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning
by: Cheng, Hanbo, et al.
Published: (2026)
by: Cheng, Hanbo, et al.
Published: (2026)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
Audio-Visual World Models: Towards Multisensory Imagination in Sight and Sound
by: Wang, Jiahua, et al.
Published: (2025)
by: Wang, Jiahua, et al.
Published: (2025)
Language-Assisted 3D Scene Understanding
by: Wu, Yanmin, et al.
Published: (2023)
by: Wu, Yanmin, et al.
Published: (2023)
Heterogeneous immune recovery after viral response through a dynamical model of feedback-driven persistence and clearance
by: Wang, Xiaoxin, et al.
Published: (2025)
by: Wang, Xiaoxin, et al.
Published: (2025)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
by: Zhu, Yinglun, et al.
Published: (2025)
by: Zhu, Yinglun, et al.
Published: (2025)
Mind-Brush: Integrating Agentic Cognitive Search and Reasoning into Image Generation
by: He, Jun, et al.
Published: (2026)
by: He, Jun, et al.
Published: (2026)
OMNIFLOW: A Physics-Grounded Multimodal Agent for Generalized Scientific Reasoning
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
Describe-then-Reason: Improving Multimodal Mathematical Reasoning through Visual Comprehension Training
by: Jia, Mengzhao, et al.
Published: (2024)
by: Jia, Mengzhao, et al.
Published: (2024)
Knowledge-enhanced Visual-Language Pretraining for Computational Pathology
by: Zhou, Xiao, et al.
Published: (2024)
by: Zhou, Xiao, et al.
Published: (2024)
A tumor-immune model of chronic myeloid leukemia with optimal immunotherapeutic protocols
by: Zhang, Haifeng, et al.
Published: (2025)
by: Zhang, Haifeng, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Superpixel Boundary Correction for Weakly-Supervised Semantic Segmentation on Histopathology Images
by: Wu, Hongyi, et al.
Published: (2025)
by: Wu, Hongyi, et al.
Published: (2025)
Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Discern Causal Links Across Modalities
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
by: Liu, Zeyu, et al.
Published: (2025)
by: Liu, Zeyu, et al.
Published: (2025)
How Likely Do LLMs with CoT Mimic Human Reasoning?
by: Bao, Guangsheng, et al.
Published: (2024)
by: Bao, Guangsheng, et al.
Published: (2024)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
by: Zhang, Huanyu, et al.
Published: (2025)
by: Zhang, Huanyu, et al.
Published: (2025)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Thinking Before Matching: A Reinforcement Reasoning Paradigm Towards General Person Re-Identification
by: Zhang, Quan, et al.
Published: (2026)
by: Zhang, Quan, et al.
Published: (2026)
Enhancing Advanced Visual Reasoning Ability of Large Language Models
by: Li, Zhiyuan, et al.
Published: (2024)
by: Li, Zhiyuan, et al.
Published: (2024)
Improving Reinforcement Learning from Human Feedback Using Contrastive Rewards
by: Shen, Wei, et al.
Published: (2024)
by: Shen, Wei, et al.
Published: (2024)
Similar Items
-
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025) -
CompilerDream: Learning a Compiler World Model for General Code Optimization
by: Deng, Chaoyi, et al.
Published: (2024) -
RLVR-World: Training World Models with Reinforcement Learning
by: Wu, Jialong, et al.
Published: (2025) -
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025) -
Domain Guidance: A Simple Transfer Approach for a Pre-trained Diffusion Model
by: Zhong, Jincheng, et al.
Published: (2025)