The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Yuwei, Yao, Yuxuan, Li, Hui, Zhu, Siyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
von: Chen, Houlun, et al.
Veröffentlicht: (2026)
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
von: Xiang, Kun, et al.
Veröffentlicht: (2024)
von: Xiang, Kun, et al.
Veröffentlicht: (2024)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
Pixel is a Barrier: Diffusion Models Are More Adversarially Robust Than We Think
von: Xue, Haotian, et al.
Veröffentlicht: (2024)
von: Xue, Haotian, et al.
Veröffentlicht: (2024)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
von: Mi, Yapeng, et al.
Veröffentlicht: (2025)
DGAE: Diffusion-Guided Autoencoder for Efficient Latent Representation Learning
von: Liu, Dongxu, et al.
Veröffentlicht: (2025)
von: Liu, Dongxu, et al.
Veröffentlicht: (2025)
Pixel-Space Post-Training of Latent Diffusion Models
von: Zhang, Christina, et al.
Veröffentlicht: (2024)
von: Zhang, Christina, et al.
Veröffentlicht: (2024)
L2P: Unlocking Latent Potential for Pixel Generation
von: Chen, Zhennan, et al.
Veröffentlicht: (2026)
von: Chen, Zhennan, et al.
Veröffentlicht: (2026)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
von: Yan, Zhonghao, et al.
Veröffentlicht: (2025)
Vision-Enhanced Time Series Forecasting via Latent Diffusion Models
von: Ruan, Weilin, et al.
Veröffentlicht: (2025)
von: Ruan, Weilin, et al.
Veröffentlicht: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
Adaptive Clinical-Aware Latent Diffusion for Multimodal Brain Image Generation and Missing Modality Imputation
von: Zhou, Rong, et al.
Veröffentlicht: (2026)
von: Zhou, Rong, et al.
Veröffentlicht: (2026)
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
von: Li, Xuchen, et al.
Veröffentlicht: (2025)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2025)
HybridStitch: Pixel and Timestep Level Model Stitching for Diffusion Acceleration
von: Sun, Desen, et al.
Veröffentlicht: (2026)
von: Sun, Desen, et al.
Veröffentlicht: (2026)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
GarmentDiffusion: 3D Garment Sewing Pattern Generation with Multimodal Diffusion Transformers
von: Li, Xinyu, et al.
Veröffentlicht: (2025)
von: Li, Xinyu, et al.
Veröffentlicht: (2025)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
Mull-Tokens: Modality-Agnostic Latent Thinking
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
von: Ray, Arijit, et al.
Veröffentlicht: (2025)
Pixelis: Reasoning in Pixels, from Seeing to Acting
von: Zhou, Yunpeng
Veröffentlicht: (2026)
von: Zhou, Yunpeng
Veröffentlicht: (2026)
Latent Action Control for Reasoning-Guided Unified Image Generation
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
von: Zhai, Fuxiang, et al.
Veröffentlicht: (2026)
EarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent Diffusion
von: Liu, Shang, et al.
Veröffentlicht: (2025)
von: Liu, Shang, et al.
Veröffentlicht: (2025)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
von: Singh, Abhishek Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Abhishek Kumar, et al.
Veröffentlicht: (2024)
GLaMM: Pixel Grounding Large Multimodal Model
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2023)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2023)
PixelBytes: Catching Unified Representation for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024)
von: Furfaro, Fabien
Veröffentlicht: (2024)
PixelBytes: Catching Unified Embedding for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024)
von: Furfaro, Fabien
Veröffentlicht: (2024)
Jailbreaks on Vision Language Model via Multimodal Reasoning
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
von: Noheria, Aarush, et al.
Veröffentlicht: (2026)
Diffusion-Guided Semantic Consistency for Multimodal Heterogeneity
von: Liu, Jing, et al.
Veröffentlicht: (2026)
von: Liu, Jing, et al.
Veröffentlicht: (2026)
IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
von: Jiang, Yankai, et al.
Veröffentlicht: (2026)
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
von: Rivkin, Dmitriy, et al.
Veröffentlicht: (2026)
von: Rivkin, Dmitriy, et al.
Veröffentlicht: (2026)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
von: Li, Yiwei, et al.
Veröffentlicht: (2026)
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
von: Xu, Haolei, et al.
Veröffentlicht: (2026)
PEAR: Pixel-aligned Expressive humAn mesh Recovery
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
von: Wu, Jiahao, et al.
Veröffentlicht: (2026)
MedScope: Incentivizing "Think with Videos" for Clinical Reasoning via Coarse-to-Fine Tool Calling
von: Li, Wenjie, et al.
Veröffentlicht: (2026)
von: Li, Wenjie, et al.
Veröffentlicht: (2026)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026) -
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
von: Chen, Houlun, et al.
Veröffentlicht: (2026) -
AtomThink: Multimodal Slow Thinking with Atomic Step Reasoning
von: Xiang, Kun, et al.
Veröffentlicht: (2024) -
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026) -
Pixel is a Barrier: Diffusion Models Are More Adversarially Robust Than We Think
von: Xue, Haotian, et al.
Veröffentlicht: (2024)