RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Wei, Xiuying, Gulcehre, Caglar |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
por: Wei, Xiuying, et al.
Publicado: (2026)
por: Wei, Xiuying, et al.
Publicado: (2026)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
por: Wei, Xiuying, et al.
Publicado: (2025)
por: Wei, Xiuying, et al.
Publicado: (2025)
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
por: Deschenaux, Justin, et al.
Publicado: (2026)
por: Deschenaux, Justin, et al.
Publicado: (2026)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
por: Deschenaux, Justin, et al.
Publicado: (2024)
por: Deschenaux, Justin, et al.
Publicado: (2024)
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
por: Terekhov, Mikhail, et al.
Publicado: (2024)
por: Terekhov, Mikhail, et al.
Publicado: (2024)
Partition Generative Modeling: Masked Modeling Without Masks
por: Deschenaux, Justin, et al.
Publicado: (2025)
por: Deschenaux, Justin, et al.
Publicado: (2025)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
por: Matrenok, Simon, et al.
Publicado: (2025)
por: Matrenok, Simon, et al.
Publicado: (2025)
The Role of Deep Learning Regularizations on Actors in Offline RL
por: Tarasov, Denis, et al.
Publicado: (2024)
por: Tarasov, Denis, et al.
Publicado: (2024)
The Diffusion Duality, Chapter II: $Ψ$-Samplers
por: Deschenaux, Justin, et al.
Publicado: (2026)
por: Deschenaux, Justin, et al.
Publicado: (2026)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
por: Orvieto, Antonio, et al.
Publicado: (2023)
por: Orvieto, Antonio, et al.
Publicado: (2023)
Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
por: Wei, Xiuying, et al.
Publicado: (2024)
por: Wei, Xiuying, et al.
Publicado: (2024)
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
por: Wei, Xiuying, et al.
Publicado: (2024)
por: Wei, Xiuying, et al.
Publicado: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
por: Bondaschi, Marco, et al.
Publicado: (2025)
por: Bondaschi, Marco, et al.
Publicado: (2025)
Loopholing Discrete Diffusion: Deterministic Bypass of the Sampling Wall
por: Jo, Mingyu, et al.
Publicado: (2025)
por: Jo, Mingyu, et al.
Publicado: (2025)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
por: Moalla, Skander, et al.
Publicado: (2024)
por: Moalla, Skander, et al.
Publicado: (2024)
HiPPO-Prophecy: State-Space Models can Provably Learn Dynamical Systems in Context
por: Joseph, Federico Arangath, et al.
Publicado: (2024)
por: Joseph, Federico Arangath, et al.
Publicado: (2024)
Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
por: Pan, Bowen, et al.
Publicado: (2024)
por: Pan, Bowen, et al.
Publicado: (2024)
Simple Hierarchical Planning with Diffusion
por: Chen, Chang, et al.
Publicado: (2024)
por: Chen, Chang, et al.
Publicado: (2024)
Control Tax: The Price of Keeping AI in Check
por: Terekhov, Mikhail, et al.
Publicado: (2025)
por: Terekhov, Mikhail, et al.
Publicado: (2025)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
por: De, Soham, et al.
Publicado: (2024)
por: De, Soham, et al.
Publicado: (2024)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
por: Surkov, Viacheslav, et al.
Publicado: (2024)
por: Surkov, Viacheslav, et al.
Publicado: (2024)
PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer
por: Chen, Chang, et al.
Publicado: (2024)
por: Chen, Chang, et al.
Publicado: (2024)
PackInfer: Compute- and I/O-Efficient Attention for Batched LLM Inference
por: Ning, Rui, et al.
Publicado: (2026)
por: Ning, Rui, et al.
Publicado: (2026)
SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
por: Zhang, Jintao, et al.
Publicado: (2025)
por: Zhang, Jintao, et al.
Publicado: (2025)
RAT: Retrieval-Augmented Transformer for Click-Through Rate Prediction
por: Li, Yushen, et al.
Publicado: (2024)
por: Li, Yushen, et al.
Publicado: (2024)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
por: Panda, Ashwinee, et al.
Publicado: (2025)
por: Panda, Ashwinee, et al.
Publicado: (2025)
Delta Attention: Fast and Accurate Sparse Attention Inference by Delta Correction
por: Willette, Jeffrey, et al.
Publicado: (2025)
por: Willette, Jeffrey, et al.
Publicado: (2025)
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
por: Karzanov, Daniil, et al.
Publicado: (2025)
por: Karzanov, Daniil, et al.
Publicado: (2025)
Neural Network Verification with PyRAT
por: Lemesle, Augustin, et al.
Publicado: (2024)
por: Lemesle, Augustin, et al.
Publicado: (2024)
Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
por: Dimitrov, Leon
Publicado: (2025)
por: Dimitrov, Leon
Publicado: (2025)
Self-Recognition in Language Models
por: Davidson, Tim R., et al.
Publicado: (2024)
por: Davidson, Tim R., et al.
Publicado: (2024)
CoreInfer: Accelerating Large Language Model Inference with Semantics-Inspired Adaptive Sparse Activation
por: Wang, Qinsi, et al.
Publicado: (2024)
por: Wang, Qinsi, et al.
Publicado: (2024)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
por: Klein, Lars, et al.
Publicado: (2024)
por: Klein, Lars, et al.
Publicado: (2024)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
por: Terekhov, Mikhail, et al.
Publicado: (2025)
por: Terekhov, Mikhail, et al.
Publicado: (2025)
Sparse Spectral Training and Inference on Euclidean and Hyperbolic Neural Networks
por: Zhao, Jialin, et al.
Publicado: (2024)
por: Zhao, Jialin, et al.
Publicado: (2024)
Training Bayesian Neural Networks with Sparse Subspace Variational Inference
por: Li, Junbo, et al.
Publicado: (2024)
por: Li, Junbo, et al.
Publicado: (2024)
RAT: Boosting Misclassification Detection Ability without Extra Data
por: Yan, Ge, et al.
Publicado: (2025)
por: Yan, Ge, et al.
Publicado: (2025)
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
por: Ye, Zihao, et al.
Publicado: (2025)
por: Ye, Zihao, et al.
Publicado: (2025)
Disentangling Dense Embeddings with Sparse Autoencoders
por: O'Neill, Charles, et al.
Publicado: (2024)
por: O'Neill, Charles, et al.
Publicado: (2024)
Inferring Latent Temporal Sparse Coordination Graph for Multi-Agent Reinforcement Learning
por: Duan, Wei, et al.
Publicado: (2024)
por: Duan, Wei, et al.
Publicado: (2024)
Ejemplares similares
-
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
por: Wei, Xiuying, et al.
Publicado: (2026) -
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
por: Wei, Xiuying, et al.
Publicado: (2025) -
BlockGen: Flexible Blockwise Sequence Modeling with Hybrid Samplers
por: Deschenaux, Justin, et al.
Publicado: (2026) -
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
por: Deschenaux, Justin, et al.
Publicado: (2024) -
In Search for Architectures and Loss Functions in Multi-Objective Reinforcement Learning
por: Terekhov, Mikhail, et al.
Publicado: (2024)