S2O: Early Stopping for Sparse Attention via Online Permutation
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yu, Liu, Songwei, Yan, Chenqian, Lin, Sheng, Ning, Beichen, Chen, Fangmin, Wang, Xing |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FoldGPT: Simple and Effective Large Language Model Compression Scheme
por: Liu, Songwei, et al.
Publicado: (2024)
por: Liu, Songwei, et al.
Publicado: (2024)
Statistical Early Stopping for Reasoning Models
por: Xie, Yangxinyu, et al.
Publicado: (2026)
por: Xie, Yangxinyu, et al.
Publicado: (2026)
ESPO: Early-Stopping Proximal Policy Optimization
por: Li, Zihang, et al.
Publicado: (2026)
por: Li, Zihang, et al.
Publicado: (2026)
Rethinking Early Stopping: Refine, Then Calibrate
por: Berta, Eugène, et al.
Publicado: (2025)
por: Berta, Eugène, et al.
Publicado: (2025)
Sparser Block-Sparse Attention via Token Permutation
por: Wang, Xinghao, et al.
Publicado: (2025)
por: Wang, Xinghao, et al.
Publicado: (2025)
An Attention-based Framework with Multistation Information for Earthquake Early Warnings
por: Huang, Yu-Ming, et al.
Publicado: (2024)
por: Huang, Yu-Ming, et al.
Publicado: (2024)
PermLLM: Learnable Channel Permutation for N:M Sparse Large Language Models
por: Zou, Lancheng, et al.
Publicado: (2025)
por: Zou, Lancheng, et al.
Publicado: (2025)
Early Stopping for Large Reasoning Models via Confidence Dynamics
por: Hosseini, Parsa, et al.
Publicado: (2026)
por: Hosseini, Parsa, et al.
Publicado: (2026)
GRAIN: Multi-Granular and Implicit Information Aggregation Graph Neural Network for Heterophilous Graphs
por: Zhao, Songwei, et al.
Publicado: (2025)
por: Zhao, Songwei, et al.
Publicado: (2025)
Stem: Rethinking Causal Information Flow in Sparse Attention
por: Niu, Lin, et al.
Publicado: (2026)
por: Niu, Lin, et al.
Publicado: (2026)
Locality Sensitive Sparse Encoding for Learning World Models Online
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
A Mathematical Theory of Top-$k$ Sparse Attention via Total Variation Distance
por: Tzachristas, Georgios, et al.
Publicado: (2025)
por: Tzachristas, Georgios, et al.
Publicado: (2025)
Don't Waste Your Time: Early Stopping Cross-Validation
por: Bergman, Edward, et al.
Publicado: (2024)
por: Bergman, Edward, et al.
Publicado: (2024)
Learning Permutation Distributions via Reflected Diffusion on Ranks
por: He, Sizhuang, et al.
Publicado: (2026)
por: He, Sizhuang, et al.
Publicado: (2026)
Towards Robust Knowledge Tracing Models via k-Sparse Attention
por: Huang, Shuyan, et al.
Publicado: (2024)
por: Huang, Shuyan, et al.
Publicado: (2024)
HISA: Efficient Hierarchical Indexing for Fine-Grained Sparse Attention
por: Xu, Yufei, et al.
Publicado: (2026)
por: Xu, Yufei, et al.
Publicado: (2026)
Learning the Optimal Stopping for Early Classification within Finite Horizons via Sequential Probability Ratio Test
por: Ebihara, Akinori F., et al.
Publicado: (2025)
por: Ebihara, Akinori F., et al.
Publicado: (2025)
vAttention: Verified Sparse Attention
por: Desai, Aditya, et al.
Publicado: (2025)
por: Desai, Aditya, et al.
Publicado: (2025)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
por: Gao, Yizhao, et al.
Publicado: (2025)
por: Gao, Yizhao, et al.
Publicado: (2025)
Intrinsically Interpretable Attention via Sparse Post-Training
por: Draye, Florent, et al.
Publicado: (2025)
por: Draye, Florent, et al.
Publicado: (2025)
Online Policy Distillation with Decision-Attention
por: Yu, Xinqiang, et al.
Publicado: (2024)
por: Yu, Xinqiang, et al.
Publicado: (2024)
Improving Sparse Autoencoder with Dynamic Attention
por: Wang, Dongsheng, et al.
Publicado: (2026)
por: Wang, Dongsheng, et al.
Publicado: (2026)
Post-Training Sparse Attention with Double Sparsity
por: Yang, Shuo, et al.
Publicado: (2024)
por: Yang, Shuo, et al.
Publicado: (2024)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
por: Yuan, Jingyang, et al.
Publicado: (2025)
por: Yuan, Jingyang, et al.
Publicado: (2025)
EntropyStop: Unsupervised Deep Outlier Detection with Loss Entropy
por: Huang, Yihong, et al.
Publicado: (2024)
por: Huang, Yihong, et al.
Publicado: (2024)
Self-Indexing KVCache: Predicting Sparse Attention from Compressed Keys
por: Yang, Xu, et al.
Publicado: (2026)
por: Yang, Xu, et al.
Publicado: (2026)
Double-P: Hierarchical Top-P Sparse Attention for Long-Context LLMs
por: Ni, Wentao, et al.
Publicado: (2026)
por: Ni, Wentao, et al.
Publicado: (2026)
Learning Unbiased Permutations via Flow Matching
por: Min, Yimeng, et al.
Publicado: (2026)
por: Min, Yimeng, et al.
Publicado: (2026)
Error Propagation Mechanisms and Compensation Strategies for Quantized Diffusion
por: Liu, Songwei, et al.
Publicado: (2025)
por: Liu, Songwei, et al.
Publicado: (2025)
Online Sparse Feature Selection in Data Streams via Differential Evolution
por: Xu, Ruiyang
Publicado: (2025)
por: Xu, Ruiyang
Publicado: (2025)
BEACON: Bayesian Optimal Stopping for Efficient LLM Sampling
por: Wan, Guangya, et al.
Publicado: (2025)
por: Wan, Guangya, et al.
Publicado: (2025)
MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inference
por: Zhou, Ruijie, et al.
Publicado: (2026)
por: Zhou, Ruijie, et al.
Publicado: (2026)
Spatial-Temporal Attention Model for Traffic State Estimation with Sparse Internet of Vehicles
por: Xue, Jianzhe, et al.
Publicado: (2024)
por: Xue, Jianzhe, et al.
Publicado: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
por: Zhu, Qianchao, et al.
Publicado: (2024)
por: Zhu, Qianchao, et al.
Publicado: (2024)
SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
por: Liu, Junnan, et al.
Publicado: (2026)
por: Liu, Junnan, et al.
Publicado: (2026)
Online Adversarial Knowledge Distillation for Graph Neural Networks
por: Wang, Can, et al.
Publicado: (2021)
por: Wang, Can, et al.
Publicado: (2021)
VSPrefill: Vertical-Slash Sparse Attention with Lightweight Indexing for Long-Context Prefilling
por: Guanzhong, Chen
Publicado: (2026)
por: Guanzhong, Chen
Publicado: (2026)
Motion-Aware Caching for Efficient Autoregressive Video Generation
por: Xu, Jing, et al.
Publicado: (2026)
por: Xu, Jing, et al.
Publicado: (2026)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
por: Deng, Yichuan, et al.
Publicado: (2024)
por: Deng, Yichuan, et al.
Publicado: (2024)
HSR-Enhanced Sparse Attention Acceleration
por: Chen, Bo, et al.
Publicado: (2024)
por: Chen, Bo, et al.
Publicado: (2024)
Ejemplares similares
-
FoldGPT: Simple and Effective Large Language Model Compression Scheme
por: Liu, Songwei, et al.
Publicado: (2024) -
Statistical Early Stopping for Reasoning Models
por: Xie, Yangxinyu, et al.
Publicado: (2026) -
ESPO: Early-Stopping Proximal Policy Optimization
por: Li, Zihang, et al.
Publicado: (2026) -
Rethinking Early Stopping: Refine, Then Calibrate
por: Berta, Eugène, et al.
Publicado: (2025) -
Sparser Block-Sparse Attention via Token Permutation
por: Wang, Xinghao, et al.
Publicado: (2025)