Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
Fuente:
arXiv
Saved in:
| Main Authors: | Garcia, Gonzalo Martin, Knaebel, Karim, Schmidt, Christian, de Geus, Daan, Hermans, Alexander, Leibe, Bastian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
by: Knaebel, Karim, et al.
Published: (2025)
by: Knaebel, Karim, et al.
Published: (2025)
SurGe: Improved Surface Geometry in Point Maps
by: Knaebel, Karim, et al.
Published: (2026)
by: Knaebel, Karim, et al.
Published: (2026)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
by: Knaebel, Karim, et al.
Published: (2023)
by: Knaebel, Karim, et al.
Published: (2023)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
by: Nekrasov, Alexey, et al.
Published: (2025)
by: Nekrasov, Alexey, et al.
Published: (2025)
How Important are Videos for Training Video LLMs?
by: Lydakis, George, et al.
Published: (2025)
by: Lydakis, George, et al.
Published: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
by: Knoche, Markus, et al.
Published: (2025)
by: Knoche, Markus, et al.
Published: (2025)
Your ViT is Secretly an Image Segmentation Model
by: Kerssies, Tommie, et al.
Published: (2025)
by: Kerssies, Tommie, et al.
Published: (2025)
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
by: Sun, Zening, et al.
Published: (2026)
by: Sun, Zening, et al.
Published: (2026)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
by: Yilmaz, Kadir, et al.
Published: (2026)
by: Yilmaz, Kadir, et al.
Published: (2026)
Representation Alignment for Just Image Transformers is not Easier than You Think
by: Shin, Jaeyo, et al.
Published: (2026)
by: Shin, Jaeyo, et al.
Published: (2026)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
by: Piekenbrinck, Jens, et al.
Published: (2025)
by: Piekenbrinck, Jens, et al.
Published: (2025)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
by: Norouzi, Narges, et al.
Published: (2026)
by: Norouzi, Narges, et al.
Published: (2026)
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
by: Schmidt, Christian, et al.
Published: (2024)
by: Schmidt, Christian, et al.
Published: (2024)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
by: Wu, Ge, et al.
Published: (2025)
by: Wu, Ge, et al.
Published: (2025)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
by: Nekrasov, Alexey, et al.
Published: (2024)
by: Nekrasov, Alexey, et al.
Published: (2024)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
by: de Geus, Daan, et al.
Published: (2024)
by: de Geus, Daan, et al.
Published: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
by: Tian, Jie, et al.
Published: (2025)
by: Tian, Jie, et al.
Published: (2025)
Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
by: Wang, Chung-Shien Brian, et al.
Published: (2025)
by: Wang, Chung-Shien Brian, et al.
Published: (2025)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
by: Cavagnero, Niccolò, et al.
Published: (2026)
by: Cavagnero, Niccolò, et al.
Published: (2026)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
by: Wienholt, Patrick, et al.
Published: (2024)
by: Wienholt, Patrick, et al.
Published: (2024)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
by: Uslu, Jan-Lucas, et al.
Published: (2024)
by: Uslu, Jan-Lucas, et al.
Published: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
by: Li, Ouxiang, et al.
Published: (2025)
by: Li, Ouxiang, et al.
Published: (2025)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
by: Kerssies, Tommie, et al.
Published: (2024)
by: Kerssies, Tommie, et al.
Published: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
by: Norouzi, Narges, et al.
Published: (2024)
by: Norouzi, Narges, et al.
Published: (2024)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
by: Yilmaz, Kadir, et al.
Published: (2023)
by: Yilmaz, Kadir, et al.
Published: (2023)
Membership Inference Attacks for Face Images Against Fine-Tuned Latent Diffusion Models
by: Holme, Lauritz Christian, et al.
Published: (2025)
by: Holme, Lauritz Christian, et al.
Published: (2025)
Fine Tuning Text-to-Image Diffusion Models for Correcting Anomalous Images
by: Yoo, Hyunwoo
Published: (2024)
by: Yoo, Hyunwoo
Published: (2024)
Point-VOS: Pointing Up Video Object Segmentation
by: Zulfikar, Idil Esen, et al.
Published: (2024)
by: Zulfikar, Idil Esen, et al.
Published: (2024)
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
by: Phunyaphibarn, Prin, et al.
Published: (2025)
by: Phunyaphibarn, Prin, et al.
Published: (2025)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
by: Englert, Brunó B., et al.
Published: (2024)
by: Englert, Brunó B., et al.
Published: (2024)
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
by: Linder, Timm, et al.
Published: (2025)
by: Linder, Timm, et al.
Published: (2025)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
by: Liu, Zheyuan, et al.
Published: (2025)
by: Liu, Zheyuan, et al.
Published: (2025)
OCCUQ: Exploring Efficient Uncertainty Quantification for 3D Occupancy Prediction
by: Heidrich, Severin, et al.
Published: (2025)
by: Heidrich, Severin, et al.
Published: (2025)
Similar Items
-
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
by: Knaebel, Karim, et al.
Published: (2025) -
SurGe: Improved Surface Geometry in Point Maps
by: Knaebel, Karim, et al.
Published: (2026) -
Point2Vec for Self-Supervised Representation Learning on Point Clouds
by: Knaebel, Karim, et al.
Published: (2023) -
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
by: Nekrasov, Alexey, et al.
Published: (2025) -
How Important are Videos for Training Video LLMs?
by: Lydakis, George, et al.
Published: (2025)