ZeroS: Zero-Sum Linear Attention for Efficient Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Jiecheng, Han, Xu, Sun, Yan, Pati, Viresh, Kim, Yubin, Somani, Siddhartha, Yang, Shihao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
by: Pati, Viresh, et al.
Published: (2026)
by: Pati, Viresh, et al.
Published: (2026)
StretchTime: Adaptive Time Series Forecasting via Symplectic Attention
by: Kim, Yubin, et al.
Published: (2026)
by: Kim, Yubin, et al.
Published: (2026)
Beyond Similarity: Temporal Operator Attention for Time Series Analysis
by: Twitty, Jevon, et al.
Published: (2026)
by: Twitty, Jevon, et al.
Published: (2026)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
by: Lu, Jiecheng, et al.
Published: (2025)
by: Lu, Jiecheng, et al.
Published: (2025)
WAVE: Weighted Autoregressive Varying Gate for Time Series Forecasting
by: Lu, Jiecheng, et al.
Published: (2024)
by: Lu, Jiecheng, et al.
Published: (2024)
CATS: Enhancing Multivariate Time Series Forecasting by Constructing Auxiliary Time Series as Exogenous Variables
by: Lu, Jiecheng, et al.
Published: (2024)
by: Lu, Jiecheng, et al.
Published: (2024)
In-context Time Series Predictor
by: Lu, Jiecheng, et al.
Published: (2024)
by: Lu, Jiecheng, et al.
Published: (2024)
HyperMLP: An Integrated Perspective for Sequence Modeling
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
Free Energy Mixer
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
by: Malmsten, Emil, et al.
Published: (2025)
by: Malmsten, Emil, et al.
Published: (2025)
MathConstraint: Automated Generation of Verified Combinatorial Reasoning Instances for LLMs
by: Pati, Viresh, et al.
Published: (2026)
by: Pati, Viresh, et al.
Published: (2026)
ARM: Refining Multivariate Forecasting with Adaptive Temporal-Contextual Learning
by: Lu, Jiecheng, et al.
Published: (2023)
by: Lu, Jiecheng, et al.
Published: (2023)
MiniZero: Comparative Analysis of AlphaZero and MuZero on Go, Othello, and Atari Games
by: Wu, Ti-Rong, et al.
Published: (2023)
by: Wu, Ti-Rong, et al.
Published: (2023)
Breaking Model Lock-in: Cost-Efficient Zero-Shot LLM Routing via a Universal Latent Space
by: Yan, Cheng, et al.
Published: (2026)
by: Yan, Cheng, et al.
Published: (2026)
Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
by: Fu, Xinghong, et al.
Published: (2026)
by: Fu, Xinghong, et al.
Published: (2026)
Regret-Guided Search Control for Efficient Learning in AlphaZero
by: Tsai, Yun-Jui, et al.
Published: (2026)
by: Tsai, Yun-Jui, et al.
Published: (2026)
ZeroShotOpt: Towards Zero-Shot Pretrained Models for Efficient Black-Box Optimization
by: Meindl, Jamison, et al.
Published: (2025)
by: Meindl, Jamison, et al.
Published: (2025)
SeqFusion: Sequential Fusion of Pre-Trained Models for Zero-Shot Time-Series Forecasting
by: Huang, Ting-Ji, et al.
Published: (2025)
by: Huang, Ting-Ji, et al.
Published: (2025)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
by: Zhang, Renrui, et al.
Published: (2023)
by: Zhang, Renrui, et al.
Published: (2023)
Zero-Shot Robustification of Zero-Shot Models
by: Adila, Dyah, et al.
Published: (2023)
by: Adila, Dyah, et al.
Published: (2023)
Molecular De Novo Design through Transformer-based Reinforcement Learning
by: Xu, Pengcheng, et al.
Published: (2023)
by: Xu, Pengcheng, et al.
Published: (2023)
Absolute Zero: Reinforced Self-play Reasoning with Zero Data
by: Zhao, Andrew, et al.
Published: (2025)
by: Zhao, Andrew, et al.
Published: (2025)
AlphaZero-Edu: Democratizing Access to AlphaZero
by: Li, Ruitong, et al.
Published: (2025)
by: Li, Ruitong, et al.
Published: (2025)
State Rank Dynamics in Linear Attention LLMs
by: Sun, Ao, et al.
Published: (2026)
by: Sun, Ao, et al.
Published: (2026)
Improving Sample Efficiency of Model-Free Algorithms for Zero-Sum Markov Games
by: Feng, Songtao, et al.
Published: (2023)
by: Feng, Songtao, et al.
Published: (2023)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
Parameter Efficient Mamba Tuning via Projector-targeted Diagonal-centric Linear Transformation
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Efficient Reinforcement Learning for Zero-Shot Coordination in Evolving Games
by: Hui, Bingyu, et al.
Published: (2025)
by: Hui, Bingyu, et al.
Published: (2025)
Communication-Efficient Byzantine-Resilient Federated Zero-Order Optimization
by: Neto, Afonso de Sá Delgado, et al.
Published: (2024)
by: Neto, Afonso de Sá Delgado, et al.
Published: (2024)
V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
by: Wang, Han, et al.
Published: (2026)
by: Wang, Han, et al.
Published: (2026)
Cooperative Open-ended Learning Framework for Zero-shot Coordination
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
OTSeg: Multi-prompt Sinkhorn Attention for Zero-Shot Semantic Segmentation
by: Kim, Kwanyoung, et al.
Published: (2024)
by: Kim, Kwanyoung, et al.
Published: (2024)
Efficient Linear Attention for Multivariate Time Series Modeling via Entropy Equality
by: Zhang, Mingtao, et al.
Published: (2025)
by: Zhang, Mingtao, et al.
Published: (2025)
Adversarial Reinforcement Learning for Offensive and Defensive Agents in a Simulated Zero-Sum Network Environment
by: Shahid, Abrar, et al.
Published: (2025)
by: Shahid, Abrar, et al.
Published: (2025)
Ister: Linear Transformer for Efficient Multivariate Time Series Forecasting
by: Cao, Fanpu, et al.
Published: (2024)
by: Cao, Fanpu, et al.
Published: (2024)
Physics-Informed Inference Time Scaling for Solving High-Dimensional PDE via Defect Correction
by: Fan, Zexi, et al.
Published: (2025)
by: Fan, Zexi, et al.
Published: (2025)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Similar Items
-
CAPS: Unifying Attention, Recurrence, and Alignment in Transformer-based Time Series Forecasting
by: Pati, Viresh, et al.
Published: (2026) -
StretchTime: Adaptive Time Series Forecasting via Symplectic Attention
by: Kim, Yubin, et al.
Published: (2026) -
Beyond Similarity: Temporal Operator Attention for Time Series Analysis
by: Twitty, Jevon, et al.
Published: (2026) -
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
by: Lu, Jiecheng, et al.
Published: (2025) -
WAVE: Weighted Autoregressive Varying Gate for Time Series Forecasting
by: Lu, Jiecheng, et al.
Published: (2024)