Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.23966 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908750087979008 |
|---|---|
| author | Zhang, Chen Bai, Yang Li, Jiahuan Gui, Anchun Wang, Keheng Liu, Feifan Wu, Guanyu Jiang, Yuwei Bu, Defei Wei, Li Jing, Haihang Tang, Hongyin Chen, Xin Huang, Xiangzhou Li, Fengcun Weng, Rongxiang Qian, Yulei Lu, Yifan Sun, Yerui Wang, Jingang Xie, Yuchen Cai, Xunliang |
| author_facet | Zhang, Chen Bai, Yang Li, Jiahuan Gui, Anchun Wang, Keheng Liu, Feifan Wu, Guanyu Jiang, Yuwei Bu, Defei Wei, Li Jing, Haihang Tang, Hongyin Chen, Xin Huang, Xiangzhou Li, Fengcun Weng, Rongxiang Qian, Yulei Lu, Yifan Sun, Yerui Wang, Jingang Xie, Yuchen Cai, Xunliang |
| contents | We introduce LongCat ZigZag Attention (LoZA), which is a sparse attention scheme designed to transform any existing full-attention models into sparse versions with rather limited compute budget. In long-context scenarios, LoZA can achieve significant speed-ups both for prefill-intensive (e.g., retrieval-augmented generation) and decode-intensive (e.g., tool-integrated reasoning) cases. Specifically, by applying LoZA to LongCat-Flash during mid-training, we serve LongCat-Flash-Exp as a long-context foundation model that can swiftly process up to 1 million tokens, enabling efficient long-term reasoning and long-horizon agentic capabilities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_23966 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Efficient Context Scaling with LongCat ZigZag Attention Zhang, Chen Bai, Yang Li, Jiahuan Gui, Anchun Wang, Keheng Liu, Feifan Wu, Guanyu Jiang, Yuwei Bu, Defei Wei, Li Jing, Haihang Tang, Hongyin Chen, Xin Huang, Xiangzhou Li, Fengcun Weng, Rongxiang Qian, Yulei Lu, Yifan Sun, Yerui Wang, Jingang Xie, Yuchen Cai, Xunliang Computation and Language Artificial Intelligence We introduce LongCat ZigZag Attention (LoZA), which is a sparse attention scheme designed to transform any existing full-attention models into sparse versions with rather limited compute budget. In long-context scenarios, LoZA can achieve significant speed-ups both for prefill-intensive (e.g., retrieval-augmented generation) and decode-intensive (e.g., tool-integrated reasoning) cases. Specifically, by applying LoZA to LongCat-Flash during mid-training, we serve LongCat-Flash-Exp as a long-context foundation model that can swiftly process up to 1 million tokens, enabling efficient long-term reasoning and long-horizon agentic capabilities. |
| title | Efficient Context Scaling with LongCat ZigZag Attention |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2512.23966 |