Saved in:
Bibliographic Details
Main Authors: Zhang, Chen, Bai, Yang, Li, Jiahuan, Gui, Anchun, Wang, Keheng, Liu, Feifan, Wu, Guanyu, Jiang, Yuwei, Bu, Defei, Wei, Li, Jing, Haihang, Tang, Hongyin, Chen, Xin, Huang, Xiangzhou, Li, Fengcun, Weng, Rongxiang, Qian, Yulei, Lu, Yifan, Sun, Yerui, Wang, Jingang, Xie, Yuchen, Cai, Xunliang
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.23966
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • We introduce LongCat ZigZag Attention (LoZA), which is a sparse attention scheme designed to transform any existing full-attention models into sparse versions with rather limited compute budget. In long-context scenarios, LoZA can achieve significant speed-ups both for prefill-intensive (e.g., retrieval-augmented generation) and decode-intensive (e.g., tool-integrated reasoning) cases. Specifically, by applying LoZA to LongCat-Flash during mid-training, we serve LongCat-Flash-Exp as a long-context foundation model that can swiftly process up to 1 million tokens, enabling efficient long-term reasoning and long-horizon agentic capabilities.