GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Phillip Y., Yoon, Taehoon, Sung, Minhyuk |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReGround: Improving Textual and Spatial Grounding at No Cost
by: Lee, Phillip Y., et al.
Published: (2024)
by: Lee, Phillip Y., et al.
Published: (2024)
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
by: Kim, Jaihoon, et al.
Published: (2025)
by: Kim, Jaihoon, et al.
Published: (2025)
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
by: Phunyaphibarn, Prin, et al.
Published: (2025)
by: Phunyaphibarn, Prin, et al.
Published: (2025)
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
by: Koo, Juil, et al.
Published: (2025)
by: Koo, Juil, et al.
Published: (2025)
Psi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models
by: Yoon, Taehoon, et al.
Published: (2025)
by: Yoon, Taehoon, et al.
Published: (2025)
GrounDiff: Diffusion-Based Ground Surface Generation from Digital Surface Models
by: Dhaouadi, Oussema, et al.
Published: (2025)
by: Dhaouadi, Oussema, et al.
Published: (2025)
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
by: Park, Mingue, et al.
Published: (2025)
by: Park, Mingue, et al.
Published: (2025)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
by: Lee, Phillip Y., et al.
Published: (2025)
by: Lee, Phillip Y., et al.
Published: (2025)
Token Warping Helps MLLMs Look from Nearby Viewpoints
by: Lee, Phillip Y., et al.
Published: (2026)
by: Lee, Phillip Y., et al.
Published: (2026)
InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion
by: Lee, Jihyun, et al.
Published: (2024)
by: Lee, Jihyun, et al.
Published: (2024)
MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos
by: Kim, Taeyeon, et al.
Published: (2026)
by: Kim, Taeyeon, et al.
Published: (2026)
Beyond the Patch: Exploring Vulnerabilities of Visuomotor Policies via Viewpoint-Consistent 3D Adversarial Object
by: Lee, Chanmi, et al.
Published: (2026)
by: Lee, Chanmi, et al.
Published: (2026)
MemBench: Memorized Image Trigger Prompt Dataset for Diffusion Models
by: Hong, Chunsan, et al.
Published: (2024)
by: Hong, Chunsan, et al.
Published: (2024)
PartSTAD: 2D-to-3D Part Segmentation Task Adaptation
by: Kim, Hyunjin, et al.
Published: (2024)
by: Kim, Hyunjin, et al.
Published: (2024)
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
by: Kim, Jaihoon, et al.
Published: (2024)
by: Kim, Jaihoon, et al.
Published: (2024)
ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
by: Min, Yunhong, et al.
Published: (2025)
by: Min, Yunhong, et al.
Published: (2025)
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
by: Yeo, Kyeongmin, et al.
Published: (2025)
by: Yeo, Kyeongmin, et al.
Published: (2025)
SALAD: Part-Level Latent Diffusion for 3D Shape Generation and Manipulation
by: Koo, Juil, et al.
Published: (2023)
by: Koo, Juil, et al.
Published: (2023)
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels
by: Guo, Erjian, et al.
Published: (2025)
by: Guo, Erjian, et al.
Published: (2025)
Proxy-Free Gaussian Splats Deformation with Splat-Based Surface Estimation
by: Kim, Jaeyeong, et al.
Published: (2025)
by: Kim, Jaeyeong, et al.
Published: (2025)
Posterior Distillation Sampling
by: Koo, Juil, et al.
Published: (2023)
by: Koo, Juil, et al.
Published: (2023)
As-Plausible-As-Possible: Plausibility-Aware Mesh Deformation Using 2D Diffusion Priors
by: Yoo, Seungwoo, et al.
Published: (2023)
by: Yoo, Seungwoo, et al.
Published: (2023)
DiTAS: Quantizing Diffusion Transformers via Enhanced Activation Smoothing
by: Dong, Zhenyuan, et al.
Published: (2024)
by: Dong, Zhenyuan, et al.
Published: (2024)
Occupancy-Based Dual Contouring
by: Hwang, Jisung, et al.
Published: (2024)
by: Hwang, Jisung, et al.
Published: (2024)
DiT4Edit: Diffusion Transformer for Image Editing
by: Feng, Kunyu, et al.
Published: (2024)
by: Feng, Kunyu, et al.
Published: (2024)
LaVin-DiT: Large Vision Diffusion Transformer
by: Wang, Zhaoqing, et al.
Published: (2024)
by: Wang, Zhaoqing, et al.
Published: (2024)
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
by: Chang, Shuning, et al.
Published: (2024)
by: Chang, Shuning, et al.
Published: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
ADP-DiT: Text-Guided Diffusion Transformer for Brain Image Generation in Alzheimer's Disease Progression
by: Lee, Juneyong, et al.
Published: (2026)
by: Lee, Juneyong, et al.
Published: (2026)
MatLat: Material Latent Space for PBR Texture Generation
by: Yeo, Kyeongmin, et al.
Published: (2025)
by: Yeo, Kyeongmin, et al.
Published: (2025)
Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
by: Seo, Jinhwan, et al.
Published: (2025)
by: Seo, Jinhwan, et al.
Published: (2025)
PTQ4DiT: Post-training Quantization for Diffusion Transformers
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
DiTFastAttn: Attention Compression for Diffusion Transformer Models
by: Yuan, Zhihang, et al.
Published: (2024)
by: Yuan, Zhihang, et al.
Published: (2024)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
by: Yang, Mengping, et al.
Published: (2026)
by: Yang, Mengping, et al.
Published: (2026)
DiTVR: Zero-Shot Diffusion Transformer for Video Restoration
by: Gao, Sicheng, et al.
Published: (2025)
by: Gao, Sicheng, et al.
Published: (2025)
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
by: Ma, Yiyang, et al.
Published: (2025)
by: Ma, Yiyang, et al.
Published: (2025)
EraserDiT: Fast Video Inpainting with Diffusion Transformer Model
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
by: Choi, Jeongsoo, et al.
Published: (2025)
by: Choi, Jeongsoo, et al.
Published: (2025)
Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
Similar Items
-
ReGround: Improving Textual and Spatial Grounding at No Cost
by: Lee, Phillip Y., et al.
Published: (2024) -
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
by: Kim, Jaihoon, et al.
Published: (2025) -
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
by: Phunyaphibarn, Prin, et al.
Published: (2025) -
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
by: Koo, Juil, et al.
Published: (2025) -
Psi-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models
by: Yoon, Taehoon, et al.
Published: (2025)