Diffusion Model Patching via Mixture-of-Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Ham, Seokil, Woo, Sangmin, Kim, Jin-Young, Go, Hyojun, Park, Byeongjun, Kim, Changick |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024)
by: Park, Byeongjun, et al.
Published: (2024)
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023)
by: Park, Byeongjun, et al.
Published: (2023)
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022)
by: Park, Byeongjun, et al.
Published: (2022)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
by: Go, Hyojun, et al.
Published: (2024)
by: Go, Hyojun, et al.
Published: (2024)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
SteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric Steering
by: Park, Byeongjun, et al.
Published: (2025)
by: Park, Byeongjun, et al.
Published: (2025)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Toward Stable World Models: Measuring and Addressing World Instability in Generative Environments
by: Kwon, Soonwoo, et al.
Published: (2025)
by: Kwon, Soonwoo, et al.
Published: (2025)
Addressing Negative Transfer in Diffusion Models
by: Go, Hyojun, et al.
Published: (2023)
by: Go, Hyojun, et al.
Published: (2023)
Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
by: Kim, Hee-Seon, et al.
Published: (2025)
by: Kim, Hee-Seon, et al.
Published: (2025)
Generating Human Motion Videos using a Cascaded Text-to-Video Framework
by: Nam, Hyelin, et al.
Published: (2025)
by: Nam, Hyelin, et al.
Published: (2025)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
by: Woo, Sangmin, et al.
Published: (2021)
by: Woo, Sangmin, et al.
Published: (2021)
Denoising Task Difficulty-based Curriculum for Training Diffusion Models
by: Kim, Jin-Young, et al.
Published: (2024)
by: Kim, Jin-Young, et al.
Published: (2024)
Spatio-Temporal Proximity-Aware Dual-Path Model for Panoramic Activity Recognition
by: Lee, Sumin, et al.
Published: (2024)
by: Lee, Sumin, et al.
Published: (2024)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
PromptDresser: Improving the Quality and Controllability of Virtual Try-On via Generative Textual Prompt and Prompt-aware Mask
by: Kim, Jeongho, et al.
Published: (2024)
by: Kim, Jeongho, et al.
Published: (2024)
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
by: Lee, Wonjun, et al.
Published: (2025)
by: Lee, Wonjun, et al.
Published: (2025)
MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completion
by: Shin, Minjung, et al.
Published: (2025)
by: Shin, Minjung, et al.
Published: (2025)
Understanding, Accelerating, and Improving MeanFlow Training
by: Kim, Jin-Young, et al.
Published: (2025)
by: Kim, Jin-Young, et al.
Published: (2025)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
by: Kim, Ju-Young, et al.
Published: (2025)
by: Kim, Ju-Young, et al.
Published: (2025)
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
Difficulty-aware Balancing Margin Loss for Long-tailed Recognition
by: Son, Minseok, et al.
Published: (2024)
by: Son, Minseok, et al.
Published: (2024)
Memory-Efficient Fine-Tuning Diffusion Transformers via Dynamic Patch Sampling and Block Skipping
by: Park, Sunghyun, et al.
Published: (2026)
by: Park, Sunghyun, et al.
Published: (2026)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
SEAL-pose: Enhancing 3D Human Pose Estimation via a Learned Loss for Structural Consistency
by: Kim, Yeonsung, et al.
Published: (2026)
by: Kim, Yeonsung, et al.
Published: (2026)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
by: Kim, Dahye, et al.
Published: (2026)
by: Kim, Dahye, et al.
Published: (2026)
Dynamic VLM-Guided Negative Prompting for Diffusion Models
by: Chang, Hoyeon, et al.
Published: (2025)
by: Chang, Hoyeon, et al.
Published: (2025)
Efficient Policy Adaptation with Contrastive Prompt Ensemble for Embodied Agents
by: Choi, Wonje, et al.
Published: (2024)
by: Choi, Wonje, et al.
Published: (2024)
Hand-object reconstruction via interaction-aware graph attention mechanism
by: Woo, Taeyun, et al.
Published: (2024)
by: Woo, Taeyun, et al.
Published: (2024)
GLYPH-SR: Can We Achieve Both High-Quality Image Super-Resolution and High-Fidelity Text Recovery via VLM-guided Latent Diffusion Model?
by: Sung, Mingyu, et al.
Published: (2025)
by: Sung, Mingyu, et al.
Published: (2025)
Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models
by: Chae, JungWoo, et al.
Published: (2025)
by: Chae, JungWoo, et al.
Published: (2025)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
by: Park, YeongHyeon, et al.
Published: (2024)
by: Park, YeongHyeon, et al.
Published: (2024)
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
by: Kim, Sunoh, et al.
Published: (2023)
by: Kim, Sunoh, et al.
Published: (2023)
REPrune: Channel Pruning via Kernel Representative Selection
by: Park, Mincheol, et al.
Published: (2024)
by: Park, Mincheol, et al.
Published: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps
by: Lee, Seoyeon, et al.
Published: (2025)
by: Lee, Seoyeon, et al.
Published: (2025)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Similar Items
-
Switch Diffusion Transformer: Synergizing Denoising Tasks with Sparse Mixture-of-Experts
by: Park, Byeongjun, et al.
Published: (2024) -
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023) -
Bridging Implicit and Explicit Geometric Transformation for Single-Image View Synthesis
by: Park, Byeongjun, et al.
Published: (2022) -
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025) -
SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis
by: Go, Hyojun, et al.
Published: (2024)