KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Youngwan, Park, Kwanyong, Cho, Yoorhim, Lee, Yong-Ju, Hwang, Sung Ju |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026)
by: Lee, Youngwan, et al.
Published: (2026)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025)
by: Lee, Youngwan, et al.
Published: (2025)
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022)
by: Hwang, Sunil, et al.
Published: (2022)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024)
by: Kim, Kangsan, et al.
Published: (2024)
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)
by: Lee, Youngwan, et al.
Published: (2024)
Identity Decoupling for Multi-Subject Personalization of Text-to-Image Models
by: Jang, Sangwon, et al.
Published: (2024)
by: Jang, Sangwon, et al.
Published: (2024)
PhaseMark: A Post-hoc, Optimization-Free Watermarking of AI-generated Images in the Latent Frequency Domain
by: Lee, Sung Ju, et al.
Published: (2026)
by: Lee, Sung Ju, et al.
Published: (2026)
Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion Models
by: Jang, Sangwon, et al.
Published: (2025)
by: Jang, Sangwon, et al.
Published: (2025)
Semantic Watermarking Reinvented: Enhancing Robustness and Generation Quality with Fourier Integrity
by: Lee, Sung Ju, et al.
Published: (2025)
by: Lee, Sung Ju, et al.
Published: (2025)
Concept-skill Transferability-based Data Selection for Large Vision-Language Models
by: Lee, Jaewoo, et al.
Published: (2024)
by: Lee, Jaewoo, et al.
Published: (2024)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
by: Kim, Kangsan, et al.
Published: (2026)
by: Kim, Kangsan, et al.
Published: (2026)
Talk in Pieces, See in Whole: Disentangling and Hierarchical Aggregating Representations for Language-based Object Detection
by: An, Sojung, et al.
Published: (2025)
by: An, Sojung, et al.
Published: (2025)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
by: Jin, Youngwan, et al.
Published: (2024)
by: Jin, Youngwan, et al.
Published: (2024)
BECoTTA: Input-dependent Online Blending of Experts for Continual Test-time Adaptation
by: Lee, Daeun, et al.
Published: (2024)
by: Lee, Daeun, et al.
Published: (2024)
ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt Tuning
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
PiLaMIM: Toward Richer Visual Representations by Integrating Pixel and Latent Masked Image Modeling
by: Lee, Junmyeong, et al.
Published: (2025)
by: Lee, Junmyeong, et al.
Published: (2025)
Improving Cone-Beam CT Image Quality with Knowledge Distillation-Enhanced Diffusion Model in Imbalanced Data Settings
by: Hwang, Joonil, et al.
Published: (2024)
by: Hwang, Joonil, et al.
Published: (2024)
FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
by: Zhang, Youyuan, et al.
Published: (2024)
by: Zhang, Youyuan, et al.
Published: (2024)
SpectraDINO: Bridging the Spectral Gap in Vision Foundation Models via Lightweight Adapters
by: Nalcakan, Yagiz, et al.
Published: (2026)
by: Nalcakan, Yagiz, et al.
Published: (2026)
PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation
by: Hwang, Injoon, et al.
Published: (2024)
by: Hwang, Injoon, et al.
Published: (2024)
GranQ: Efficient Channel-wise Quantization via Vectorized Pre-Scaling for Zero-Shot QAT
by: Hong, Inpyo, et al.
Published: (2025)
by: Hong, Inpyo, et al.
Published: (2025)
A Training-free Sub-quadratic Cost Transformer Model Serving Framework With Hierarchically Pruned Attention
by: Lee, Heejun, et al.
Published: (2024)
by: Lee, Heejun, et al.
Published: (2024)
Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
by: Kim, Chanyoung, et al.
Published: (2025)
by: Kim, Chanyoung, et al.
Published: (2025)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
by: Han, Woojung, et al.
Published: (2025)
by: Han, Woojung, et al.
Published: (2025)
VideoPatchCore: An Effective Method to Memorize Normality for Video Anomaly Detection
by: Ahn, Sunghyun, et al.
Published: (2024)
by: Ahn, Sunghyun, et al.
Published: (2024)
Memory-Efficient Personalization of Text-to-Image Diffusion Models via Selective Optimization Strategies
by: Choi, Seokeon, et al.
Published: (2025)
by: Choi, Seokeon, et al.
Published: (2025)
Exposing Text-Image Inconsistency Using Diffusion Models
by: Huang, Mingzhen, et al.
Published: (2024)
by: Huang, Mingzhen, et al.
Published: (2024)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
by: Lee, Uichan, et al.
Published: (2026)
by: Lee, Uichan, et al.
Published: (2026)
Towards Label-Efficient Human Matting: A Simple Baseline for Weakly Semi-Supervised Trimap-Free Human Matting
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Text-Aware Image Restoration with Diffusion Models
by: Min, Jaewon, et al.
Published: (2025)
by: Min, Jaewon, et al.
Published: (2025)
Flatfish Lesion Detection Based on Part Segmentation Approach and Lesion Image Generation
by: Hwang, Seo-Bin, et al.
Published: (2024)
by: Hwang, Seo-Bin, et al.
Published: (2024)
Rethinking Saliency-Guided Weakly-Supervised Semantic Segmentation
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
Towards Spatially Consistent Image Generation: On Incorporating Intrinsic Scene Properties into Diffusion Models
by: Lee, Hyundo, et al.
Published: (2025)
by: Lee, Hyundo, et al.
Published: (2025)
Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
by: Seo, Hoigi, et al.
Published: (2026)
by: Seo, Hoigi, et al.
Published: (2026)
VM-DDPM: Vision Mamba Diffusion for Medical Image Synthesis
by: Ju, Zhihan, et al.
Published: (2024)
by: Ju, Zhihan, et al.
Published: (2024)
Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
by: Kang, Seongjae, et al.
Published: (2025)
by: Kang, Seongjae, et al.
Published: (2025)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
by: Jin, Youngwan, et al.
Published: (2026)
by: Jin, Youngwan, et al.
Published: (2026)
CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple Images
by: Shin, Jisu, et al.
Published: (2024)
by: Shin, Jisu, et al.
Published: (2024)
Designing Extremely Memory-Efficient CNNs for On-device Vision Tasks
by: Lee, Jaewook, et al.
Published: (2024)
by: Lee, Jaewook, et al.
Published: (2024)
Similar Items
-
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2026) -
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
by: Lee, Youngwan, et al.
Published: (2025) -
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens
by: Hwang, Sunil, et al.
Published: (2022) -
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024) -
Visualizing the loss landscape of Self-supervised Vision Transformer
by: Lee, Youngwan, et al.
Published: (2024)