Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Steifer, Tomasz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences
von: Tumma, Neehal, et al.
Veröffentlicht: (2026)
von: Tumma, Neehal, et al.
Veröffentlicht: (2026)
Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection
von: Lee, Tae-Gyun, et al.
Veröffentlicht: (2026)
von: Lee, Tae-Gyun, et al.
Veröffentlicht: (2026)
Simple online learning with consistent oracle
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2023)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2023)
A completely uniform transformer for parity
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2025)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2025)
Computable universal online learning
von: Kalociński, Dariusz, et al.
Veröffentlicht: (2025)
von: Kalociński, Dariusz, et al.
Veröffentlicht: (2025)
Parity, Sensitivity, and Transformers
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
Ehrenfeucht-Haussler Rank and Chain of Thought
von: Barceló, Pablo, et al.
Veröffentlicht: (2025)
von: Barceló, Pablo, et al.
Veröffentlicht: (2025)
Effective Littlestone Dimension
von: Rose, Valentino Delle, et al.
Veröffentlicht: (2024)
von: Rose, Valentino Delle, et al.
Veröffentlicht: (2024)
Optimal bounds for dissatisfaction in perpetual voting
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2024)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2024)
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention
von: Zhou, Chenyu, et al.
Veröffentlicht: (2026)
von: Zhou, Chenyu, et al.
Veröffentlicht: (2026)
CausalEvolve: Towards Open-Ended Discovery with Causal Scratchpad
von: Chen, Yongqiang, et al.
Veröffentlicht: (2026)
von: Chen, Yongqiang, et al.
Veröffentlicht: (2026)
Strassen Attention, Split VC Dimension and Compositionality in Transformers
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2025)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2025)
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2024)
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2024)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
von: Harel, Itamar, et al.
Veröffentlicht: (2024)
von: Harel, Itamar, et al.
Veröffentlicht: (2024)
Provably Learning Attention with Queries
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2026)
von: Bhattamishra, Satwik, et al.
Veröffentlicht: (2026)
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models
von: Zheng, Lin, et al.
Veröffentlicht: (2026)
von: Zheng, Lin, et al.
Veröffentlicht: (2026)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
von: Ye, Xiaowei, et al.
Veröffentlicht: (2026)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
von: Willette, Jeffrey, et al.
Veröffentlicht: (2025)
von: Willette, Jeffrey, et al.
Veröffentlicht: (2025)
Provably tuning the ElasticNet across instances
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2022)
von: Balcan, Maria-Florina, et al.
Veröffentlicht: (2022)
Provable Generalization in Overparameterized Neural Nets
von: Dhingra, Aviral
Veröffentlicht: (2025)
von: Dhingra, Aviral
Veröffentlicht: (2025)
Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMs
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
von: Liu, Hongyi, et al.
Veröffentlicht: (2025)
Task-Aware Calibration: Provably Optimal Decoding in LLMs
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
von: Tomov, Tim, et al.
Veröffentlicht: (2026)
KQ-SVD: Compressing the KV Cache with Provable Guarantees on Attention Fidelity
von: Lesens, Damien, et al.
Veröffentlicht: (2025)
von: Lesens, Damien, et al.
Veröffentlicht: (2025)
Delta Attention Residuals
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
von: Luo, Cheng, et al.
Veröffentlicht: (2026)
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
von: Li, Yichen, et al.
Veröffentlicht: (2024)
von: Li, Yichen, et al.
Veröffentlicht: (2024)
Attention with Trained Embeddings Provably Selects Important Tokens
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
von: Wu, Diyuan, et al.
Veröffentlicht: (2025)
Questioning the Coverage-Length Metric in Conformal Prediction: When Shorter Intervals Are Not Better
von: Min, Yizhou, et al.
Veröffentlicht: (2026)
von: Min, Yizhou, et al.
Veröffentlicht: (2026)
Minimalist Softmax Attention Provably Learns Constrained Boolean Functions
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
von: Cai, Changxiao, et al.
Veröffentlicht: (2026)
von: Cai, Changxiao, et al.
Veröffentlicht: (2026)
G-Net: A Provably Easy Construction of High-Accuracy Random Binary Neural Networks
von: Aghasi, Alireza, et al.
Veröffentlicht: (2025)
von: Aghasi, Alireza, et al.
Veröffentlicht: (2025)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
von: Ran-Milo, Yuval
Veröffentlicht: (2026)
Provable Differentially Private Computation of the Cross-Attention Mechanism
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
von: Ke, Yekun, et al.
Veröffentlicht: (2024)
Hybrid Transfer Reinforcement Learning: Provable Sample Efficiency from Shifted-Dynamics Data
von: Qu, Chengrui, et al.
Veröffentlicht: (2024)
von: Qu, Chengrui, et al.
Veröffentlicht: (2024)
Less Effort, Shorter Proofs: Reinforcement Learning for Security Protocol Analysis in Tamarin
von: Cosler, Matthias, et al.
Veröffentlicht: (2026)
von: Cosler, Matthias, et al.
Veröffentlicht: (2026)
C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness
von: Kang, Yu, et al.
Veröffentlicht: (2024)
von: Kang, Yu, et al.
Veröffentlicht: (2024)
Invertible ResNets for Inverse Imaging Problems: Competitive Performance with Provable Regularization Properties
von: Arndt, Clemens, et al.
Veröffentlicht: (2024)
von: Arndt, Clemens, et al.
Veröffentlicht: (2024)
Transformers Provably Learn Sparse Token Selection While Fully-Connected Nets Cannot
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
CAT-Net: A Cross-Attention Tone Network for Cross-Subject EEG-EMG Fusion Tone Decoding
von: Zhuang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhuang, Yifan, et al.
Veröffentlicht: (2025)
Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression
von: Tian, Ye, et al.
Veröffentlicht: (2026)
von: Tian, Ye, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences
von: Tumma, Neehal, et al.
Veröffentlicht: (2026) -
Patched-DeltaNet: Token-Level Event-Driven Memory for Linear-Time Anomaly Detection
von: Lee, Tae-Gyun, et al.
Veröffentlicht: (2026) -
Simple online learning with consistent oracle
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2023) -
A completely uniform transformer for parity
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2025) -
Computable universal online learning
von: Kalociński, Dariusz, et al.
Veröffentlicht: (2025)