The Crystal Ball Hypothesis in diffusion models: Anticipating object positions from initial noise
Fuente:
arXiv
Saved in:
| Main Authors: | Ban, Yuanhao, Wang, Ruochen, Zhou, Tianyi, Gong, Boqing, Hsieh, Cho-Jui, Cheng, Minhao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
by: Ban, Yuanhao, et al.
Published: (2024)
by: Ban, Yuanhao, et al.
Published: (2024)
MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
by: Li, Sen, et al.
Published: (2024)
by: Li, Sen, et al.
Published: (2024)
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
by: Zhou, Hengguang, et al.
Published: (2025)
by: Zhou, Hengguang, et al.
Published: (2025)
On Discrete Prompt Optimization for Diffusion Models
by: Wang, Ruochen, et al.
Published: (2024)
by: Wang, Ruochen, et al.
Published: (2024)
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
One-Forcing: Towards Stable One-Step Autoregressive Video Generation
by: Feng, Jiaqi, et al.
Published: (2026)
by: Feng, Jiaqi, et al.
Published: (2026)
Mitigating Bias in Dataset Distillation
by: Cui, Justin, et al.
Published: (2024)
by: Cui, Justin, et al.
Published: (2024)
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding
by: Xiong, Yuanhao, et al.
Published: (2023)
by: Xiong, Yuanhao, et al.
Published: (2023)
IRIS: Intrinsic Reward Image Synthesis
by: Chen, Yihang, et al.
Published: (2025)
by: Chen, Yihang, et al.
Published: (2025)
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
by: Kao, Kuei-Chun, et al.
Published: (2026)
by: Kao, Kuei-Chun, et al.
Published: (2026)
Concepts or Skills? Rethinking Instruction Selection for Multi-modal Models
by: Bai, Andrew, et al.
Published: (2025)
by: Bai, Andrew, et al.
Published: (2025)
LoL: Longer than Longer, Scaling Video Generation to Hour
by: Cui, Justin, et al.
Published: (2026)
by: Cui, Justin, et al.
Published: (2026)
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
by: Cui, Justin, et al.
Published: (2025)
by: Cui, Justin, et al.
Published: (2025)
Understanding Reward Hacking in Text-to-Image Reinforcement Learning
by: Hong, Yunqi, et al.
Published: (2026)
by: Hong, Yunqi, et al.
Published: (2026)
Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
by: Kao, Kuei-Chun, et al.
Published: (2025)
by: Kao, Kuei-Chun, et al.
Published: (2025)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
by: Gong, Dayoung, et al.
Published: (2024)
by: Gong, Dayoung, et al.
Published: (2024)
Culture in Action: Evaluating Text-to-Image Models through Social Activities
by: Malakouti, Sina, et al.
Published: (2025)
by: Malakouti, Sina, et al.
Published: (2025)
Attention to Neural Plagiarism: Diffusion Models Can Plagiarize Your Copyrighted Images!
by: Zou, Zihang, et al.
Published: (2026)
by: Zou, Zihang, et al.
Published: (2026)
HypDAE: Hyperbolic Diffusion Autoencoders for Hierarchical Few-shot Image Generation
by: Li, Lingxiao, et al.
Published: (2024)
by: Li, Lingxiao, et al.
Published: (2024)
Embedding Space Selection for Detecting Memorization and Fingerprinting in Generative Models
by: He, Jack, et al.
Published: (2024)
by: He, Jack, et al.
Published: (2024)
Blue noise for diffusion models
by: Huang, Xingchang, et al.
Published: (2024)
by: Huang, Xingchang, et al.
Published: (2024)
Continual Adapter Tuning with Semantic Shift Compensation for Class-Incremental Learning
by: Zhou, Qinhao, et al.
Published: (2024)
by: Zhou, Qinhao, et al.
Published: (2024)
Adversarial Examples Detection with Bayesian Neural Network
by: Li, Yao, et al.
Published: (2021)
by: Li, Yao, et al.
Published: (2021)
Large Language Models are Interpretable Learners
by: Wang, Ruochen, et al.
Published: (2024)
by: Wang, Ruochen, et al.
Published: (2024)
Short-term Object Interaction Anticipation with Disentangled Object Detection @ Ego4D Short Term Object Interaction Anticipation Challenge
by: Cho, Hyunjin, et al.
Published: (2024)
by: Cho, Hyunjin, et al.
Published: (2024)
EGGS: Edge Guided Gaussian Splatting for Radiance Fields
by: Gong, Yuanhao
Published: (2024)
by: Gong, Yuanhao
Published: (2024)
BlurBall: Joint Ball and Motion Blur Estimation for Table Tennis Ball Tracking
by: Gossard, Thomas, et al.
Published: (2025)
by: Gossard, Thomas, et al.
Published: (2025)
Do text-free diffusion models learn discriminative visual representations?
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
by: Mukhopadhyay, Soumik, et al.
Published: (2023)
Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation
by: Luo, Yuanhao, et al.
Published: (2026)
by: Luo, Yuanhao, et al.
Published: (2026)
DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
by: Li, Xirui, et al.
Published: (2024)
by: Li, Xirui, et al.
Published: (2024)
VideoAds for Fast-Paced Video Understanding
by: Zhang, Zheyuan, et al.
Published: (2025)
by: Zhang, Zheyuan, et al.
Published: (2025)
LDM-Morph: Latent diffusion model guided deformable image registration
by: Wu, Jiong, et al.
Published: (2024)
by: Wu, Jiong, et al.
Published: (2024)
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
by: Labadia, Lorenzo Mur, et al.
Published: (2026)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
by: Tan, Yuwen, et al.
Published: (2025)
by: Tan, Yuwen, et al.
Published: (2025)
Edge-preserving noise for diffusion models
by: Vandersanden, Jente, et al.
Published: (2024)
by: Vandersanden, Jente, et al.
Published: (2024)
JIGMARK: A Black-Box Approach for Enhancing Image Watermarks against Diffusion Model Edits
by: Pan, Minzhou, et al.
Published: (2024)
by: Pan, Minzhou, et al.
Published: (2024)
Text is All You Need for Vision-Language Model Jailbreaking
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
Echo-DND: A dual noise diffusion model for robust and precise left ventricle segmentation in echocardiography
by: Rahman, Abdur, et al.
Published: (2025)
by: Rahman, Abdur, et al.
Published: (2025)
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
by: Mur-Labadia, Lorenzo, et al.
Published: (2024)
Similar Items
-
Understanding the Impact of Negative Prompts: When and How Do They Take Effect?
by: Ban, Yuanhao, et al.
Published: (2024) -
MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
by: Li, Sen, et al.
Published: (2024) -
R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model
by: Zhou, Hengguang, et al.
Published: (2025) -
On Discrete Prompt Optimization for Diffusion Models
by: Wang, Ruochen, et al.
Published: (2024) -
MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
by: Li, Xirui, et al.
Published: (2024)