DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | He, Zefeng, Qu, Xiaoye, Li, Yafu, Zhu, Tong, Huang, Siyuan, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
GEMS: Agent-Native Multimodal Generation with Memory and Skills
by: He, Zefeng, et al.
Published: (2026)
by: He, Zefeng, et al.
Published: (2026)
Spotlight on Token Perception for Multimodal Reinforcement Learning
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
by: Huang, Siyuan, et al.
Published: (2026)
by: Huang, Siyuan, et al.
Published: (2026)
CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling
by: Zhang, Jihai, et al.
Published: (2024)
by: Zhang, Jihai, et al.
Published: (2024)
SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
by: Shen, Chuming, et al.
Published: (2025)
by: Shen, Chuming, et al.
Published: (2025)
Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning
by: Liao, Xinyao, et al.
Published: (2025)
by: Liao, Xinyao, et al.
Published: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
SketchThinker-R1: Towards Efficient Sketch-Style Reasoning in Large Multimodal Models
by: Zhang, Ruiyang, et al.
Published: (2026)
by: Zhang, Ruiyang, et al.
Published: (2026)
Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
by: Chen, Shuang, et al.
Published: (2025)
by: Chen, Shuang, et al.
Published: (2025)
Towards Stabilized and Efficient Diffusion Transformers through Long-Skip-Connections with Spectral Constraints
by: Chen, Guanjie, et al.
Published: (2024)
by: Chen, Guanjie, et al.
Published: (2024)
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
by: Pei, Baoqi, et al.
Published: (2025)
by: Pei, Baoqi, et al.
Published: (2025)
DiffBIR: Towards Blind Image Restoration with Generative Diffusion Prior
by: Lin, Xinqi, et al.
Published: (2023)
by: Lin, Xinqi, et al.
Published: (2023)
Look, Compare, Decide: Alleviating Hallucination in Large Vision-Language Models via Multi-View Multi-Path Reasoning
by: Qu, Xiaoye, et al.
Published: (2024)
by: Qu, Xiaoye, et al.
Published: (2024)
Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model
by: Li, Jing, et al.
Published: (2023)
by: Li, Jing, et al.
Published: (2023)
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
ReactDiff: Fundamental Multiple Appropriate Facial Reaction Diffusion Model
by: Cheng, Luo, et al.
Published: (2025)
by: Cheng, Luo, et al.
Published: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Unified Thinker: A General Reasoning Modular Core for Image Generation
by: Zhou, Sashuai, et al.
Published: (2026)
by: Zhou, Sashuai, et al.
Published: (2026)
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning
by: Li, Xueheng, et al.
Published: (2026)
by: Li, Xueheng, et al.
Published: (2026)
OneThinker: All-in-one Reasoning Model for Image and Video
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
by: Cheng, Zixu, et al.
Published: (2026)
by: Cheng, Zixu, et al.
Published: (2026)
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
by: Wu, Jianzong, et al.
Published: (2024)
by: Wu, Jianzong, et al.
Published: (2024)
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
by: Su, Zhaochen, et al.
Published: (2025)
by: Su, Zhaochen, et al.
Published: (2025)
DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
by: Han, Feng, et al.
Published: (2025)
by: Han, Feng, et al.
Published: (2025)
Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
by: Chen, Guanjie, et al.
Published: (2025)
by: Chen, Guanjie, et al.
Published: (2025)
StgcDiff: Spatial-Temporal Graph Condition Diffusion for Sign Language Transition Generation
by: He, Jiashu, et al.
Published: (2025)
by: He, Jiashu, et al.
Published: (2025)
DiffFAS: Face Anti-Spoofing via Generative Diffusion Models
by: Ge, Xinxu, et al.
Published: (2024)
by: Ge, Xinxu, et al.
Published: (2024)
Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension
by: Xu, Haoran, et al.
Published: (2026)
by: Xu, Haoran, et al.
Published: (2026)
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
by: Jiang, Qing, et al.
Published: (2025)
by: Jiang, Qing, et al.
Published: (2025)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
by: Liao, Xinyao, et al.
Published: (2026)
by: Liao, Xinyao, et al.
Published: (2026)
DiffMAC: Diffusion Manifold Hallucination Correction for High Generalization Blind Face Restoration
by: Gao, Nan, et al.
Published: (2024)
by: Gao, Nan, et al.
Published: (2024)
ViThinker: Active Vision-Language Reasoning via Dynamic Perceptual Querying
by: You, Weihang, et al.
Published: (2026)
by: You, Weihang, et al.
Published: (2026)
DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
by: Yu, Wenwen, et al.
Published: (2025)
by: Yu, Wenwen, et al.
Published: (2025)
CycleDiff: Cycle Diffusion Models for Unpaired Image-to-image Translation
by: Zou, Shilong, et al.
Published: (2025)
by: Zou, Shilong, et al.
Published: (2025)
DiffInk: Glyph- and Style-Aware Latent Diffusion Transformer for Text to Online Handwriting Generation
by: Pan, Wei, et al.
Published: (2025)
by: Pan, Wei, et al.
Published: (2025)
VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
by: Wang, Qunzhong, et al.
Published: (2025)
by: Wang, Qunzhong, et al.
Published: (2025)
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
by: Ruan, Jiacheng, et al.
Published: (2026)
by: Ruan, Jiacheng, et al.
Published: (2026)
Similar Items
-
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
by: He, Zefeng, et al.
Published: (2025) -
GEMS: Agent-Native Multimodal Generation with Memory and Skills
by: He, Zefeng, et al.
Published: (2026) -
Spotlight on Token Perception for Multimodal Reinforcement Learning
by: Huang, Siyuan, et al.
Published: (2025) -
VideoSSR: Video Self-Supervised Reinforcement Learning
by: He, Zefeng, et al.
Published: (2025) -
Persistent Visual Memory: Sustaining Perception for Deep Generation in LVLMs
by: Huang, Siyuan, et al.
Published: (2026)