I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mi, Zhenxing, Wang, Kuan-Chieh, Qian, Guocheng, Ye, Hanrong, Liu, Runtao, Tulyakov, Sergey, Aberman, Kfir, Xu, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
von: Wang, Kuan-Chieh, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chieh, et al.
Veröffentlicht: (2024)
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
von: Dalva, Yusuf, et al.
Veröffentlicht: (2025)
von: Dalva, Yusuf, et al.
Veröffentlicht: (2025)
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)
Therefore I am. I Think
von: Esakkiraja, Esakkivel, et al.
Veröffentlicht: (2026)
von: Esakkiraja, Esakkivel, et al.
Veröffentlicht: (2026)
Interpreting the Weight Space of Customized Diffusion Models
von: Dravid, Amil, et al.
Veröffentlicht: (2024)
von: Dravid, Amil, et al.
Veröffentlicht: (2024)
MyVLM: Personalizing VLMs for User-Specific Queries
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
Orthogonal Adaptation for Modular Customization of Diffusion Models
von: Po, Ryan, et al.
Veröffentlicht: (2023)
von: Po, Ryan, et al.
Veröffentlicht: (2023)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
von: Patashnik, Or, et al.
Veröffentlicht: (2025)
von: Patashnik, Or, et al.
Veröffentlicht: (2025)
Preventing Shortcuts in Adapter Training via Providing the Shortcuts
von: Goyal, Anujraaj Argo, et al.
Veröffentlicht: (2025)
von: Goyal, Anujraaj Argo, et al.
Veröffentlicht: (2025)
Diffusion-DRF: Free, Rich, and Differentiable Reward for Video Diffusion Fine-Tuning
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
von: Wang, Yifan, et al.
Veröffentlicht: (2026)
AToM: Amortized Text-to-Mesh using 2D Diffusion
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
I Think, Therefore I Hallucinate: Minds, Machines, and the Art of Being Wrong
von: Barros, Sebastian
Veröffentlicht: (2025)
von: Barros, Sebastian
Veröffentlicht: (2025)
I Teach, Therefore I Hope
von: Gelacio, Joseph
Veröffentlicht: (2026)
von: Gelacio, Joseph
Veröffentlicht: (2026)
I Think Therefore I Am: Building Ethical AI Through Relational Training
von: Willoughby, Samuel James
Veröffentlicht: (2026)
von: Willoughby, Samuel James
Veröffentlicht: (2026)
I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench
von: Li, Yuan, et al.
Veröffentlicht: (2024)
von: Li, Yuan, et al.
Veröffentlicht: (2024)
Diffusion Priors for Dynamic View Synthesis from Monocular Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Dynamic Concepts Personalization from Single Videos
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Chapter 9 I Think, Therefore IR? Psychology, Biology and the Notion of Praxis
von: Davis, James W.
Veröffentlicht: (2022)
von: Davis, James W.
Veröffentlicht: (2022)
Continuous Control of Editing Models via Adaptive-Origin Guidance
von: Wolf, Alon, et al.
Veröffentlicht: (2026)
von: Wolf, Alon, et al.
Veröffentlicht: (2026)
Scaling Group Inference for Diverse and High-Quality Generation
von: Parmar, Gaurav, et al.
Veröffentlicht: (2025)
von: Parmar, Gaurav, et al.
Veröffentlicht: (2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
SPAD : Spatially Aware Multiview Diffusers
von: Kant, Yash, et al.
Veröffentlicht: (2024)
von: Kant, Yash, et al.
Veröffentlicht: (2024)
Visual Personalization Turing Test
von: Abdal, Rameen, et al.
Veröffentlicht: (2026)
von: Abdal, Rameen, et al.
Veröffentlicht: (2026)
Efficient Training with Denoised Neural Weights
von: Gong, Yifan, et al.
Veröffentlicht: (2024)
von: Gong, Yifan, et al.
Veröffentlicht: (2024)
LeC$^2$O-NeRF: Learning Continuous and Compact Large-Scale Occupancy for Urban Scenes
von: Mi, Zhenxing, et al.
Veröffentlicht: (2024)
von: Mi, Zhenxing, et al.
Veröffentlicht: (2024)
I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
von: Kharchenko, Julia, et al.
Veröffentlicht: (2025)
von: Kharchenko, Julia, et al.
Veröffentlicht: (2025)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
Be Yourself: Bounded Attention for Multi-Subject Text-to-Image Generation
von: Dahary, Omer, et al.
Veröffentlicht: (2024)
von: Dahary, Omer, et al.
Veröffentlicht: (2024)
I See, Therefore I Do: Estimating Causal Effects for Image Treatments
von: Thorat, Abhinav, et al.
Veröffentlicht: (2024)
von: Thorat, Abhinav, et al.
Veröffentlicht: (2024)
I Move Therefore I Learn: Experience-Based Traversability in Outdoor Robotics
von: de Miguel, Miguel Ángel, et al.
Veröffentlicht: (2025)
von: de Miguel, Miguel Ángel, et al.
Veröffentlicht: (2025)
Object-level Visual Prompts for Compositional Image Generation
von: Parmar, Gaurav, et al.
Veröffentlicht: (2025)
von: Parmar, Gaurav, et al.
Veröffentlicht: (2025)
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
von: Kim, Keuntae, et al.
Veröffentlicht: (2026)
3D PixBrush: Image-Guided Local Texture Synthesis
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
von: Decatur, Dale, et al.
Veröffentlicht: (2025)
ArtifactLens: Hundreds of Labels Are Enough for Artifact Detection with VLMs
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
Guitar Tone Morphing by Diffusion-based Model
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
E$^{2}$GAN: Efficient Training of Efficient GANs for Image-to-Image Translation
von: Gong, Yifan, et al.
Veröffentlicht: (2024)
von: Gong, Yifan, et al.
Veröffentlicht: (2024)
Diffuse Thinking: Exploring Diffusion Language Models as Efficient Thought Proposers for Reasoning
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
von: Shao, Chenyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
von: Wang, Kuan-Chieh, et al.
Veröffentlicht: (2024) -
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
von: Qian, Guocheng, et al.
Veröffentlicht: (2024) -
Canvas-to-Image: Compositional Image Generation with Multimodal Controls
von: Dalva, Yusuf, et al.
Veröffentlicht: (2025) -
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025) -
DiffusionMTL: Learning Multi-Task Denoising Diffusion Model from Partially Annotated Data
von: Ye, Hanrong, et al.
Veröffentlicht: (2024)