MADI: Masking-Augmented Diffusion with Inference-Time Scaling for Visual Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kadambi, Shreya, Garrepalli, Risheek, Borse, Shubhankar, Hyatt, Munawar, Porikli, Fatih |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024)
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
FouRA: Fourier Low Rank Adaptation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024)
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
OCAI: Improving Optical Flow Estimation by Occlusion and Consistency Aware Interpolation
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
von: Jeong, Jisoo, et al.
Veröffentlicht: (2024)
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
von: Farhadzadeh, Farzad, et al.
Veröffentlicht: (2025)
Personalized OVSS: Understanding Personal Concept in Open-Vocabulary Semantic Segmentation
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
von: Park, Sunghyun, et al.
Veröffentlicht: (2025)
PosSAM: Panoptic Open-vocabulary Segment Anything
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
Clockwork Diffusion: Efficient Generation With Model-Step Distillation
von: Habibian, Amirhossein, et al.
Veröffentlicht: (2023)
von: Habibian, Amirhossein, et al.
Veröffentlicht: (2023)
SubZero: Composing Subject, Style, and Action via Zero-Shot Personalization
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
MAMo: Leveraging Memory and Attention for Monocular Video Depth Estimation
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2023)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2023)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
SciFlow: Empowering Lightweight Optical Flow Models with Self-Cleaning Iterations
von: Lin, Jamie Menjay, et al.
Veröffentlicht: (2024)
von: Lin, Jamie Menjay, et al.
Veröffentlicht: (2024)
Attention Guided Alignment in Efficient Vision-Language Models
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)
Generalized Contrastive Learning for Universal Multimodal Retrieval
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
Ar2Can: An Architect and an Artist Leveraging a Canvas for Multi-Human Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025)
A Comparative analysis of Layer-wise Representational Capacity in AR and Diffusion LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
von: Goel, Raghavv, et al.
Veröffentlicht: (2026)
FutureDepth: Learning to Predict the Future Improves Video Depth Estimation
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2024)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2024)
Spiffy: Multiplying Diffusion LLM Acceleration via Lossless Speculative Decoding
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
von: Agrawal, Sudhanshu, et al.
Veröffentlicht: (2025)
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
Hidden Bias in the Machine: Stereotypes in Text-to-Image Models
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
von: Porikli, Sedat, et al.
Veröffentlicht: (2025)
DICE: Discrete Inversion Enabling Controllable Editing for Multinomial Diffusion and Masked Generative Models
von: He, Xiaoxiao, et al.
Veröffentlicht: (2024)
von: He, Xiaoxiao, et al.
Veröffentlicht: (2024)
Distilling Multi-modal Large Language Models for Autonomous Driving
von: Hegde, Deepti, et al.
Veröffentlicht: (2025)
von: Hegde, Deepti, et al.
Veröffentlicht: (2025)
MoViE: Mobile Diffusion for Video Editing
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
Textualize Visual Prompt for Image Editing via Diffusion Bridge
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2025)
Hollowed Net for On-Device Personalization of Text-to-Image Diffusion Models
von: Cho, Wonguk, et al.
Veröffentlicht: (2024)
von: Cho, Wonguk, et al.
Veröffentlicht: (2024)
Performance Plateaus in Inference-Time Scaling for Text-to-Image Diffusion Without External Models
von: Choi, Changhyun, et al.
Veröffentlicht: (2025)
von: Choi, Changhyun, et al.
Veröffentlicht: (2025)
Click2Mask: Local Editing with Dynamic Mask Generation
von: Regev, Omer, et al.
Veröffentlicht: (2024)
von: Regev, Omer, et al.
Veröffentlicht: (2024)
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
von: Cho, Janghoon, et al.
Veröffentlicht: (2025)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
Masked Generative Transformer Is What You Need for Image Editing
von: Chow, Wei, et al.
Veröffentlicht: (2026)
von: Chow, Wei, et al.
Veröffentlicht: (2026)
Inference-Time Scaling of Diffusion Models for Infrared Data Generation
von: Horstmann, Kai A., et al.
Veröffentlicht: (2025)
von: Horstmann, Kai A., et al.
Veröffentlicht: (2025)
DuoLoRA : Cycle-consistent and Rank-disentangled Content-Style Personalization
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
von: Lee, Jungsoo, et al.
Veröffentlicht: (2025)
Two-Step Data Augmentation for Masked Face Detection and Recognition: Turning Fake Masks to Real
von: Yang, Yan, et al.
Veröffentlicht: (2025)
von: Yang, Yan, et al.
Veröffentlicht: (2025)
Unified Control for Inference-Time Guidance of Denoising Diffusion Models
von: Goyal, Maurya, et al.
Veröffentlicht: (2025)
von: Goyal, Maurya, et al.
Veröffentlicht: (2025)
Unified Concept Editing in Diffusion Models
von: Gandikota, Rohit, et al.
Veröffentlicht: (2023)
von: Gandikota, Rohit, et al.
Veröffentlicht: (2023)
Tripartite Weight-Space Ensemble for Few-Shot Class-Incremental Learning
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
von: Lee, Juntae, et al.
Veröffentlicht: (2025)
Mask and Restore: Blind Backdoor Defense at Test Time with Masked Autoencoder
von: Sun, Tao, et al.
Veröffentlicht: (2023)
von: Sun, Tao, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
DDIL: Diversity Enhancing Diffusion Distillation With Imitation Learning
von: Garrepalli, Risheek, et al.
Veröffentlicht: (2024) -
MultiHuman-Testbench: Benchmarking Image Generation for Multiple Humans
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025) -
FouRA: Fourier Low Rank Adaptation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2024) -
Resolving the Identity Crisis in Text-to-Image Generation
von: Borse, Shubhankar, et al.
Veröffentlicht: (2025) -
Do-Undo Bench: Reversibility for Action Understanding in Image Generation
von: Mahajan, Shweta, et al.
Veröffentlicht: (2025)