Control-DINO: Feature Space Conditioning for Controllable Image-to-Video Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Dominici, Edoardo A., Deixelberger, Thomas, Vardis, Konstantinos, Steinberger, Markus |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation
by: Ainetter, Stefan, et al.
Published: (2026)
by: Ainetter, Stefan, et al.
Published: (2026)
DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
by: Dominici, Edoardo Alberto, et al.
Published: (2025)
by: Dominici, Edoardo Alberto, et al.
Published: (2025)
SOF: Sorted Opacity Fields for Fast Unbounded Surface Reconstruction
by: Radl, Lukas, et al.
Published: (2025)
by: Radl, Lukas, et al.
Published: (2025)
Back to the Features: DINO as a Foundation for Video World Models
by: Baldassarre, Federico, et al.
Published: (2025)
by: Baldassarre, Federico, et al.
Published: (2025)
Multi-task Image Restoration Guided By Robust DINO Features
by: Lin, Xin, et al.
Published: (2023)
by: Lin, Xin, et al.
Published: (2023)
Pi-GS: Sparse-View Gaussian Splatting with Dense π^3 Initialization
by: Hofer, Manuel, et al.
Published: (2026)
by: Hofer, Manuel, et al.
Published: (2026)
DINO-Tracker: Taming DINO for Self-Supervised Point Tracking in a Single Video
by: Tumanyan, Narek, et al.
Published: (2024)
by: Tumanyan, Narek, et al.
Published: (2024)
ControlMambaIR: Conditional Controls with State-Space Model for Image Restoration
by: Yang, Cheng, et al.
Published: (2025)
by: Yang, Cheng, et al.
Published: (2025)
Local Conditional Controlling for Text-to-Image Diffusion Models
by: Zhao, Yibo, et al.
Published: (2023)
by: Zhao, Yibo, et al.
Published: (2023)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
by: Yang, Sicheng, et al.
Published: (2025)
by: Yang, Sicheng, et al.
Published: (2025)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
Feature-Conditioned Cascaded Video Diffusion Models for Precise Echocardiogram Synthesis
by: Reynaud, Hadrien, et al.
Published: (2023)
by: Reynaud, Hadrien, et al.
Published: (2023)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
ACD: Direct Conditional Control for Video Diffusion Models via Attention Supervision
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
DINO-AD: Unsupervised Anomaly Detection with Frozen DINO-V3 Features
by: Huo, Jiayu, et al.
Published: (2026)
by: Huo, Jiayu, et al.
Published: (2026)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
by: Ansari, Muhammad Musab
Published: (2025)
by: Ansari, Muhammad Musab
Published: (2025)
Graph Conditioned Diffusion for Controllable Histopathology Image Generation
by: Cechnicka, Sarah, et al.
Published: (2025)
by: Cechnicka, Sarah, et al.
Published: (2025)
Controlling Space and Time with Diffusion Models
by: Watson, Daniel, et al.
Published: (2024)
by: Watson, Daniel, et al.
Published: (2024)
Active Sampling for Ultra-Low-Bit-Rate Video Compression via Conditional Controlled Diffusion
by: Javadi, Amirhosein, et al.
Published: (2026)
by: Javadi, Amirhosein, et al.
Published: (2026)
HeartBeat: Towards Controllable Echocardiography Video Synthesis with Multimodal Conditions-Guided Diffusion Models
by: Zhou, Xinrui, et al.
Published: (2024)
by: Zhou, Xinrui, et al.
Published: (2024)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation
by: Wang, Cong, et al.
Published: (2024)
by: Wang, Cong, et al.
Published: (2024)
LAENeRF: Local Appearance Editing for Neural Radiance Fields
by: Radl, Lukas, et al.
Published: (2023)
by: Radl, Lukas, et al.
Published: (2023)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model
by: Zheng, Guangcong, et al.
Published: (2024)
by: Zheng, Guangcong, et al.
Published: (2024)
RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature Guidance
by: Fan, Jiaojiao, et al.
Published: (2024)
by: Fan, Jiaojiao, et al.
Published: (2024)
Volumetric Conditioning Module to Control Pretrained Diffusion Models for 3D Medical Images
by: Ahn, Suhyun, et al.
Published: (2024)
by: Ahn, Suhyun, et al.
Published: (2024)
PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
by: Fu, Weifu, et al.
Published: (2026)
by: Fu, Weifu, et al.
Published: (2026)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
by: Guo, Hao, et al.
Published: (2024)
by: Guo, Hao, et al.
Published: (2024)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
In-Context Audio Control of Video Diffusion Transformers
by: Liu, Wenze, et al.
Published: (2025)
by: Liu, Wenze, et al.
Published: (2025)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
by: Gong, Ziren, et al.
Published: (2025)
by: Gong, Ziren, et al.
Published: (2025)
Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
by: Liu, Shilong, et al.
Published: (2023)
by: Liu, Shilong, et al.
Published: (2023)
DC-ControlNet: Decoupling Inter- and Intra-Element Conditions in Image Generation with Diffusion Models
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
Readout Guidance: Learning Control from Diffusion Features
by: Luo, Grace, et al.
Published: (2023)
by: Luo, Grace, et al.
Published: (2023)
DIVE: Taming DINO for Subject-Driven Video Editing
by: Huang, Yi, et al.
Published: (2024)
by: Huang, Yi, et al.
Published: (2024)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
by: Liang, Zhuonan, et al.
Published: (2026)
by: Liang, Zhuonan, et al.
Published: (2026)
LaMD: Latent Motion Diffusion for Image-Conditional Video Generation
by: Hu, Yaosi, et al.
Published: (2023)
by: Hu, Yaosi, et al.
Published: (2023)
Studying Image Diffusion Features for Zero-Shot Video Object Segmentation
by: Delatolas, Thanos, et al.
Published: (2025)
by: Delatolas, Thanos, et al.
Published: (2025)
Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions
by: Zhang, David Junhao, et al.
Published: (2024)
by: Zhang, David Junhao, et al.
Published: (2024)
Similar Items
-
Scene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene Generation
by: Ainetter, Stefan, et al.
Published: (2026) -
DreamAnywhere: Object-Centric Panoramic 3D Scene Generation
by: Dominici, Edoardo Alberto, et al.
Published: (2025) -
SOF: Sorted Opacity Fields for Fast Unbounded Surface Reconstruction
by: Radl, Lukas, et al.
Published: (2025) -
Back to the Features: DINO as a Foundation for Video World Models
by: Baldassarre, Federico, et al.
Published: (2025) -
Multi-task Image Restoration Guided By Robust DINO Features
by: Lin, Xin, et al.
Published: (2023)