Accelerating Controllable Generation via Hybrid-grained Cache

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Lin, Ben, Huixia, Wang, Shuo, Lu, Jinda, Qiu, Junxiang, Tang, Shengeng, Hao, Yanbin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909901970735104
author Liu, Lin
Ben, Huixia
Wang, Shuo
Lu, Jinda
Qiu, Junxiang
Tang, Shengeng
Hao, Yanbin
author_facet Liu, Lin
Ben, Huixia
Wang, Shuo
Lu, Jinda
Qiu, Junxiang
Tang, Shengeng
Hao, Yanbin
contents Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T to 6.70T), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%.
format Preprint
id arxiv_https___arxiv_org_abs_2511_11031
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Accelerating Controllable Generation via Hybrid-grained Cache
Liu, Lin
Ben, Huixia
Wang, Shuo
Lu, Jinda
Qiu, Junxiang
Tang, Shengeng
Hao, Yanbin
Computer Vision and Pattern Recognition
Multimedia
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T to 6.70T), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%.
title Accelerating Controllable Generation via Hybrid-grained Cache
topic Computer Vision and Pattern Recognition
Multimedia
url https://arxiv.org/abs/2511.11031