S2O: Early Stopping for Sparse Attention via Online Permutation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yu, Liu, Songwei, Yan, Chenqian, Lin, Sheng, Ning, Beichen, Chen, Fangmin, Wang, Xing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FoldGPT: Simple and Effective Large Language Model Compression Scheme
di: Liu, Songwei, et al.
Pubblicazione: (2024)
di: Liu, Songwei, et al.
Pubblicazione: (2024)
Statistical Early Stopping for Reasoning Models
di: Xie, Yangxinyu, et al.
Pubblicazione: (2026)
di: Xie, Yangxinyu, et al.
Pubblicazione: (2026)
ESPO: Early-Stopping Proximal Policy Optimization
di: Li, Zihang, et al.
Pubblicazione: (2026)
di: Li, Zihang, et al.
Pubblicazione: (2026)
Rethinking Early Stopping: Refine, Then Calibrate
di: Berta, Eugène, et al.
Pubblicazione: (2025)
di: Berta, Eugène, et al.
Pubblicazione: (2025)
Sparser Block-Sparse Attention via Token Permutation
di: Wang, Xinghao, et al.
Pubblicazione: (2025)
di: Wang, Xinghao, et al.
Pubblicazione: (2025)
An Attention-based Framework with Multistation Information for Earthquake Early Warnings
di: Huang, Yu-Ming, et al.
Pubblicazione: (2024)
di: Huang, Yu-Ming, et al.
Pubblicazione: (2024)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
di: Zou, Lancheng, et al.
Pubblicazione: (2025)
di: Zou, Lancheng, et al.
Pubblicazione: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
di: Hosseini, Parsa, et al.
Pubblicazione: (2026)
GRAIN: Multi-Granular and Implicit Information Aggregation Graph Neural Network for Heterophilous Graphs
di: Zhao, Songwei, et al.
Pubblicazione: (2025)
di: Zhao, Songwei, et al.
Pubblicazione: (2025)
Stem: Rethinking Causal Information Flow in Sparse Attention
di: Niu, Lin, et al.
Pubblicazione: (2026)
di: Niu, Lin, et al.
Pubblicazione: (2026)
Locality Sensitive Sparse Encoding for Learning World Models Online
di: Liu, Zichen, et al.
Pubblicazione: (2024)
di: Liu, Zichen, et al.
Pubblicazione: (2024)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
di: Tzachristas, Georgios, et al.
Pubblicazione: (2025)
di: Tzachristas, Georgios, et al.
Pubblicazione: (2025)
Don't Waste Your Time: Early Stopping Cross-Validation
di: Bergman, Edward, et al.
Pubblicazione: (2024)
di: Bergman, Edward, et al.
Pubblicazione: (2024)
Learning Permutation Distributions via Reflected Diffusion on Ranks
di: He, Sizhuang, et al.
Pubblicazione: (2026)
di: He, Sizhuang, et al.
Pubblicazione: (2026)
Towards Robust Knowledge Tracing Models via k-Sparse Attention
di: Huang, Shuyan, et al.
Pubblicazione: (2024)
di: Huang, Shuyan, et al.
Pubblicazione: (2024)
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
di: Xu, Yufei, et al.
Pubblicazione: (2026)
di: Xu, Yufei, et al.
Pubblicazione: (2026)
Learning the Optimal Stopping for Early Classification within Finite Horizons via Sequential Probability Ratio Test
di: Ebihara, Akinori F., et al.
Pubblicazione: (2025)
di: Ebihara, Akinori F., et al.
Pubblicazione: (2025)
vAttention: Verified Sparse Attention
di: Desai, Aditya, et al.
Pubblicazione: (2025)
di: Desai, Aditya, et al.
Pubblicazione: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
di: Gao, Yizhao, et al.
Pubblicazione: (2025)
di: Gao, Yizhao, et al.
Pubblicazione: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
di: Draye, Florent, et al.
Pubblicazione: (2025)
di: Draye, Florent, et al.
Pubblicazione: (2025)
Online Policy Distillation with Decision-Attention
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
di: Yu, Xinqiang, et al.
Pubblicazione: (2024)
Improving Sparse Autoencoder with Dynamic Attention
di: Wang, Dongsheng, et al.
Pubblicazione: (2026)
di: Wang, Dongsheng, et al.
Pubblicazione: (2026)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
di: Yuan, Jingyang, et al.
Pubblicazione: (2025)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
di: Huang, Yihong, et al.
Pubblicazione: (2024)
di: Huang, Yihong, et al.
Pubblicazione: (2024)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
di: Yang, Xu, et al.
Pubblicazione: (2026)
di: Yang, Xu, et al.
Pubblicazione: (2026)
Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
di: Ni, Wentao, et al.
Pubblicazione: (2026)
di: Ni, Wentao, et al.
Pubblicazione: (2026)
Learning Unbiased Permutations via Flow Matching
di: Min, Yimeng, et al.
Pubblicazione: (2026)
di: Min, Yimeng, et al.
Pubblicazione: (2026)
Error Propagation Mechanisms and Compensation Strategies for Quantized Diffusion
di: Liu, Songwei, et al.
Pubblicazione: (2025)
di: Liu, Songwei, et al.
Pubblicazione: (2025)
Online Sparse Feature Selection in Data Streams via Differential Evolution
di: Xu, Ruiyang
Pubblicazione: (2025)
di: Xu, Ruiyang
Pubblicazione: (2025)
BEACON: Bayesian Optimal Stopping for Efficient LLM Sampling
di: Wan, Guangya, et al.
Pubblicazione: (2025)
di: Wan, Guangya, et al.
Pubblicazione: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
di: Zhou, Ruijie, et al.
Pubblicazione: (2026)
Spatial-Temporal Attention Model for Traffic State Estimation with Sparse Internet of Vehicles
di: Xue, Jianzhe, et al.
Pubblicazione: (2024)
di: Xue, Jianzhe, et al.
Pubblicazione: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
di: Zhu, Qianchao, et al.
Pubblicazione: (2024)
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
di: Liu, Junnan, et al.
Pubblicazione: (2026)
di: Liu, Junnan, et al.
Pubblicazione: (2026)
Online Adversarial Knowledge Distillation for Graph Neural Networks
di: Wang, Can, et al.
Pubblicazione: (2021)
di: Wang, Can, et al.
Pubblicazione: (2021)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
di: Guanzhong, Chen
Pubblicazione: (2026)
di: Guanzhong, Chen
Pubblicazione: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
di: Deng, Yichuan, et al.
Pubblicazione: (2024)
HSR-Enhanced Sparse Attention Acceleration
di: Chen, Bo, et al.
Pubblicazione: (2024)
di: Chen, Bo, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FoldGPT: Simple and Effective Large Language Model Compression Scheme
di: Liu, Songwei, et al.
Pubblicazione: (2024) -
Statistical Early Stopping for Reasoning Models
di: Xie, Yangxinyu, et al.
Pubblicazione: (2026) -
ESPO: Early-Stopping Proximal Policy Optimization
di: Li, Zihang, et al.
Pubblicazione: (2026) -
Rethinking Early Stopping: Refine, Then Calibrate
di: Berta, Eugène, et al.
Pubblicazione: (2025) -
Sparser Block-Sparse Attention via Token Permutation
di: Wang, Xinghao, et al.
Pubblicazione: (2025)