YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wu, You, Chen, Ziheng, Zhang, Yizhen, Wu, Haoyi, Yu, Chengting, Xu, Yuchi, Su, Wenbo, Zheng, Bo, Tu, Kewei
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917409661648896
author Wu, You
Chen, Ziheng
Zhang, Yizhen
Wu, Haoyi
Yu, Chengting
Xu, Yuchi
Su, Wenbo
Zheng, Bo
Tu, Kewei
author_facet Wu, You
Chen, Ziheng
Zhang, Yizhen
Wu, Haoyi
Yu, Chengting
Xu, Yuchi
Su, Wenbo
Zheng, Bo
Tu, Kewei
contents Cross-layer key-value (KV) compression has been found to be effective in efficient inference of large language models (LLMs). Although they reduce the memory consumption of the KV cache, such methods usually introduce non-negligible performance degradation. In this work, we aim to enhance the performance of YOCO, a cross-layer KV compression method that shares the KVs of the middle layer with the top-half layers. We propose YOCO++, an enhanced YOCO that incorporates a weighted residual connection between the KVs of each bottom-half layer and the bottom layer. Compared to YOCO, YOCO++ increases model capacity while maintaining the same training and inference efficiency. Our experiments show that YOCO++ achieves state-of-the-art performance among the cross-layer KV compression methods at a 50% KV cache compression rate, outperforming the standard Transformer.
format Preprint
id arxiv_https___arxiv_org_abs_2604_13556
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
Wu, You
Chen, Ziheng
Zhang, Yizhen
Wu, Haoyi
Yu, Chengting
Xu, Yuchi
Su, Wenbo
Zheng, Bo
Tu, Kewei
Computation and Language
Cross-layer key-value (KV) compression has been found to be effective in efficient inference of large language models (LLMs). Although they reduce the memory consumption of the KV cache, such methods usually introduce non-negligible performance degradation. In this work, we aim to enhance the performance of YOCO, a cross-layer KV compression method that shares the KVs of the middle layer with the top-half layers. We propose YOCO++, an enhanced YOCO that incorporates a weighted residual connection between the KVs of each bottom-half layer and the bottom layer. Compared to YOCO, YOCO++ increases model capacity while maintaining the same training and inference efficiency. Our experiments show that YOCO++ achieves state-of-the-art performance among the cross-layer KV compression methods at a 50% KV cache compression rate, outperforming the standard Transformer.
title YOCO++: Enhancing YOCO with KV Residual Connections for Efficient LLM Inference
topic Computation and Language
url https://arxiv.org/abs/2604.13556