Recurrent Memory-Augmented Transformers with Chunked Attention for Long-Context Language Modeling
Fuente:
arXiv
Saved in:
| Main Author: | Kashyap, Ankit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Hidden Attention of Mamba Models
by: Ali, Ameen, et al.
Published: (2024)
by: Ali, Ameen, et al.
Published: (2024)
Time-to-Injury Forecasting in Elite Female Football: A DeepHit Survival Approach
by: Catterall, Victoria, et al.
Published: (2026)
by: Catterall, Victoria, et al.
Published: (2026)
A Prescriptive Framework for Determining Optimal Days for Short-Term Traffic Counts
by: Mukwaya, Arthur, et al.
Published: (2025)
by: Mukwaya, Arthur, et al.
Published: (2025)
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
by: Pratyush, Spandan
Published: (2026)
by: Pratyush, Spandan
Published: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)
by: Atad, Ido Andrew, et al.
Published: (2026)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
by: Moghadasi, Mahdi Naser, et al.
Published: (2026)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
by: Huang, Yunpeng, et al.
Published: (2023)
by: Huang, Yunpeng, et al.
Published: (2023)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
CBR -- Boosting Adaptive Classification By Retrieval of Encrypted Network Traffic with Out-of-distribution
by: Lukach, Amir, et al.
Published: (2024)
by: Lukach, Amir, et al.
Published: (2024)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
by: Schesch, Benedikt, et al.
Published: (2026)
by: Schesch, Benedikt, et al.
Published: (2026)
TRIM: Achieving Extreme Sparsity with Targeted Row-wise Iterative Metric-driven Pruning
by: Beck, Florentin, et al.
Published: (2025)
by: Beck, Florentin, et al.
Published: (2025)
The Architecture of Errors: From Universal Impossibility to Patch-Local LLM Reliability
by: Arbuzov, Mikhail L., et al.
Published: (2026)
by: Arbuzov, Mikhail L., et al.
Published: (2026)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
by: Salfati, Samuel
Published: (2026)
by: Salfati, Samuel
Published: (2026)
Random Heterogeneous Neurochaos Learning Architecture for Data Classification
by: S, Remya Ajai A, et al.
Published: (2024)
by: S, Remya Ajai A, et al.
Published: (2024)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
by: Heyman, Alex, et al.
Published: (2025)
by: Heyman, Alex, et al.
Published: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
by: Bakish, Yarden, et al.
Published: (2025)
by: Bakish, Yarden, et al.
Published: (2025)
Enhanced and Efficient Reasoning in Large Learning Models
by: Valiant, Leslie G.
Published: (2026)
by: Valiant, Leslie G.
Published: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
by: Katraouras, Thomas, et al.
Published: (2026)
by: Katraouras, Thomas, et al.
Published: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
by: Rozonoyer, Benjamin, et al.
Published: (2026)
by: Rozonoyer, Benjamin, et al.
Published: (2026)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
by: Bianchessi, Arthur S., et al.
Published: (2025)
by: Bianchessi, Arthur S., et al.
Published: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
by: Saxena, Udit
Published: (2025)
by: Saxena, Udit
Published: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
by: Ye, Hua, et al.
Published: (2025)
by: Ye, Hua, et al.
Published: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
by: Mi, Zhendong, et al.
Published: (2025)
by: Mi, Zhendong, et al.
Published: (2025)
QuAnTS: Question Answering on Time Series
by: Divo, Felix, et al.
Published: (2025)
by: Divo, Felix, et al.
Published: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
by: Easley, Eric, et al.
Published: (2026)
by: Easley, Eric, et al.
Published: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
by: McCann, Jordan F.
Published: (2026)
by: McCann, Jordan F.
Published: (2026)
Super Apriel: One Checkpoint, Many Speeds
by: Labs, SLAM, et al.
Published: (2026)
by: Labs, SLAM, et al.
Published: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
by: Rossi, Maximillian, et al.
Published: (2026)
by: Rossi, Maximillian, et al.
Published: (2026)
DRO-InstructZero: Distributionally Robust Prompt Optimization for Large Language Models
by: Li, Yangyang
Published: (2025)
by: Li, Yangyang
Published: (2025)
Language Models Are Implicitly Continuous
by: Marro, Samuele, et al.
Published: (2025)
by: Marro, Samuele, et al.
Published: (2025)
SPHERICAL KV: Angle-Domain Attention and Rate-Distortion Retention for Efficient Long-Context Inference
by: Chauhan, Anay, et al.
Published: (2026)
by: Chauhan, Anay, et al.
Published: (2026)
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
by: Li, Zehao, et al.
Published: (2026)
by: Li, Zehao, et al.
Published: (2026)
Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention Formulation
by: Zimerman, Itamar, et al.
Published: (2024)
by: Zimerman, Itamar, et al.
Published: (2024)
Combining Language and Topic Models for Hierarchical Text Classification
by: Toit, Jaco du, et al.
Published: (2025)
by: Toit, Jaco du, et al.
Published: (2025)
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning
by: Das, Susmit
Published: (2026)
by: Das, Susmit
Published: (2026)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
by: Zhu, Jiajun, et al.
Published: (2025)
by: Zhu, Jiajun, et al.
Published: (2025)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
by: Anderson, Samuel Cyrenius
Published: (2026)
by: Anderson, Samuel Cyrenius
Published: (2026)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
by: Szilvasy, Gergely, et al.
Published: (2026)
by: Szilvasy, Gergely, et al.
Published: (2026)
Similar Items
-
The Hidden Attention of Mamba Models
by: Ali, Ameen, et al.
Published: (2024) -
Time-to-Injury Forecasting in Elite Female Football: A DeepHit Survival Approach
by: Catterall, Victoria, et al.
Published: (2026) -
A Prescriptive Framework for Determining Optimal Days for Short-Term Traffic Counts
by: Mukwaya, Arthur, et al.
Published: (2025) -
Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers
by: Pratyush, Spandan
Published: (2026) -
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
by: Atad, Ido Andrew, et al.
Published: (2026)