RNNs are not Transformers (Yet): The Key Bottleneck on In-context Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wen, Kaiyue, Dang, Xingyu, Lyu, Kaifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Power of Power Law: Asymmetry Enables Compositional Reasoning
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)
On Efficiently Representing Regular Languages as RNNs
von: Svete, Anej, et al.
Veröffentlicht: (2024)
von: Svete, Anej, et al.
Veröffentlicht: (2024)
Why Are Linear RNNs More Parallelizable?
von: Merrill, William, et al.
Veröffentlicht: (2026)
von: Merrill, William, et al.
Veröffentlicht: (2026)
How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM Pretraining
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
Learning State-Tracking from Code Using Linear RNNs
von: Siems, Julien, et al.
Veröffentlicht: (2026)
von: Siems, Julien, et al.
Veröffentlicht: (2026)
PaTH Attention: Position Encoding via Accumulating Householder Transformations
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
von: Yang, Songlin, et al.
Veröffentlicht: (2025)
Curse of High Dimensionality Issue in Transformer for Long-context Modeling
von: Zhang, Shuhai, et al.
Veröffentlicht: (2025)
von: Zhang, Shuhai, et al.
Veröffentlicht: (2025)
Transformers are Universal In-context Learners
von: Furuya, Takashi, et al.
Veröffentlicht: (2024)
von: Furuya, Takashi, et al.
Veröffentlicht: (2024)
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
von: Grazzi, Riccardo, et al.
Veröffentlicht: (2024)
Enhancing Transformer RNNs with Multiple Temporal Perspectives
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
von: Dumitru, Razvan-Gabriel, et al.
Veröffentlicht: (2024)
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
von: Gu, Xinran, et al.
Veröffentlicht: (2025)
von: Gu, Xinran, et al.
Veröffentlicht: (2025)
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
von: Tang, Kexian, et al.
Veröffentlicht: (2026)
von: Tang, Kexian, et al.
Veröffentlicht: (2026)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
Breaking the Attention Bottleneck
von: Hilsenbek, Kalle
Veröffentlicht: (2024)
von: Hilsenbek, Kalle
Veröffentlicht: (2024)
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
von: Siems, Julien, et al.
Veröffentlicht: (2025)
von: Siems, Julien, et al.
Veröffentlicht: (2025)
Bottleneck-Minimal Indexing for Generative Document Retrieval
von: Du, Xin, et al.
Veröffentlicht: (2024)
von: Du, Xin, et al.
Veröffentlicht: (2024)
Linking In-context Learning in Transformers to Human Episodic Memory
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024)
LLMs for Bayesian Optimization in Scientific Domains: Are We There Yet?
von: Gupta, Rushil, et al.
Veröffentlicht: (2025)
von: Gupta, Rushil, et al.
Veröffentlicht: (2025)
Efficient Stagewise Pretraining via Progressive Subnetworks
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
von: Panigrahi, Abhishek, et al.
Veröffentlicht: (2024)
Concept Bottleneck Large Language Models
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
von: Sun, Chung-En, et al.
Veröffentlicht: (2024)
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
A Multi-Power Law for Loss Curve Prediction Across Learning Rate Schedules
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
von: Luo, Kairong, et al.
Veröffentlicht: (2025)
MUDDFormer: Breaking Residual Bottlenecks in Transformers via Multiway Dynamic Dense Connections
von: Xiao, Da, et al.
Veröffentlicht: (2025)
von: Xiao, Da, et al.
Veröffentlicht: (2025)
Robust Training of Vector Quantized Bottleneck Models
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
von: Łańcucki, Adrian, et al.
Veröffentlicht: (2020)
Retrieval Enhanced Feedback via In-context Neural Error-book
von: Hyun, Jongyeop, et al.
Veröffentlicht: (2025)
von: Hyun, Jongyeop, et al.
Veröffentlicht: (2025)
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention
von: Brandon, William, et al.
Veröffentlicht: (2024)
von: Brandon, William, et al.
Veröffentlicht: (2024)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
von: Lyu, Kaifeng, et al.
Veröffentlicht: (2024)
von: Lyu, Kaifeng, et al.
Veröffentlicht: (2024)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models
von: Lv, Ang, et al.
Veröffentlicht: (2024)
von: Lv, Ang, et al.
Veröffentlicht: (2024)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
von: Qiu, Yifu, et al.
Veröffentlicht: (2025)
Benchmarked Yet Not Measured -- Generative AI Should be Evaluated Against Real-World Utility
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
von: Mondal, Ishani, et al.
Veröffentlicht: (2026)
A Systematic Review of Key Retrieval-Augmented Generation (RAG) Systems: Progress, Gaps, and Future Directions
von: Oche, Agada Joseph, et al.
Veröffentlicht: (2025)
von: Oche, Agada Joseph, et al.
Veröffentlicht: (2025)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
von: Wu, Songhao, et al.
Veröffentlicht: (2025)
Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence
von: Xiao, Liu
Veröffentlicht: (2026)
von: Xiao, Liu
Veröffentlicht: (2026)
Graph of Records: Boosting Retrieval Augmented Generation for Long-context Summarization with Graphs
von: Zhang, Haozhen, et al.
Veröffentlicht: (2024)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Power of Power Law: Asymmetry Enables Compositional Reasoning
von: Wang, Zixuan, et al.
Veröffentlicht: (2026) -
Does Transformer Interpretability Transfer to RNNs?
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024) -
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025) -
From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
von: Wen, Kaiyue, et al.
Veröffentlicht: (2024) -
Comba: Improving Bilinear RNNs with Closed-loop Control
von: Hu, Jiaxi, et al.
Veröffentlicht: (2025)