Scaling Transformer to 1M tokens and beyond with RMT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bulatov, Aydar, Kuratov, Yuri, Kapushev, Yermek, Burtsev, Mikhail S. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Associative Recurrent Memory Transformer
von: Rodkin, Ivan, et al.
Veröffentlicht: (2024)
von: Rodkin, Ivan, et al.
Veröffentlicht: (2024)
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
von: Chepurova, Alla, et al.
Veröffentlicht: (2025)
von: Chepurova, Alla, et al.
Veröffentlicht: (2025)
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
von: Kuratov, Yuri, et al.
Veröffentlicht: (2025)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2025)
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)
SRMT: Shared Memory for Multi-agent Lifelong Pathfinding
von: Sagirova, Alsu, et al.
Veröffentlicht: (2025)
von: Sagirova, Alsu, et al.
Veröffentlicht: (2025)
Beyond Memorization: Extending Reasoning Depth with Recurrence, Memory and Test-Time Compute Scaling
von: Rodkin, Ivan, et al.
Veröffentlicht: (2025)
von: Rodkin, Ivan, et al.
Veröffentlicht: (2025)
Looking beyond the next token
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
von: Thankaraj, Abitha, et al.
Veröffentlicht: (2025)
Long Input Benchmark for Russian Analysis
von: Churin, Igor, et al.
Veröffentlicht: (2024)
von: Churin, Igor, et al.
Veröffentlicht: (2024)
Limitations of Normalization in Attention Mechanism
von: Mudarisov, Timur, et al.
Veröffentlicht: (2025)
von: Mudarisov, Timur, et al.
Veröffentlicht: (2025)
Learning Elementary Cellular Automata with Transformers
von: Burtsev, Mikhail
Veröffentlicht: (2024)
von: Burtsev, Mikhail
Veröffentlicht: (2024)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
The pitfalls of next-token prediction
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
von: Bachmann, Gregor, et al.
Veröffentlicht: (2024)
Shaping capabilities with token-level data filtering
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
von: Rathi, Neil, et al.
Veröffentlicht: (2026)
Essential-Web v1.0: 24T tokens of organized web data
von: AI, Essential, et al.
Veröffentlicht: (2025)
von: AI, Essential, et al.
Veröffentlicht: (2025)
Interpretable Next-token Prediction via the Generalized Induction Head
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
Language models are better than humans at next-token prediction
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
von: Shlegeris, Buck, et al.
Veröffentlicht: (2022)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
All or None: Identifiable Linear Properties of Next-token Predictors in Language Modeling
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
von: Marconato, Emanuele, et al.
Veröffentlicht: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Is Sanskrit the most token-efficient language? A quantitative study using GPT, Gemini, and SentencePiece
von: Kumar, Anshul
Veröffentlicht: (2026)
von: Kumar, Anshul
Veröffentlicht: (2026)
Diagonal Batching Unlocks Parallelism in Recurrent Memory Transformers for Long Contexts
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
von: Sivtsov, Danil, et al.
Veröffentlicht: (2025)
Mixture of Chapters: Scaling Learnt Memory in Transformers
von: Tibrewal, Tasmay Pankaj, et al.
Veröffentlicht: (2026)
von: Tibrewal, Tasmay Pankaj, et al.
Veröffentlicht: (2026)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
von: Yang, Chiwun
Veröffentlicht: (2025)
von: Yang, Chiwun
Veröffentlicht: (2025)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
Revisiting Tree Search for LLMs: Gumbel and Sequential Halving for Budget-Scalable Reasoning
von: Ugadiarov, Leonid, et al.
Veröffentlicht: (2026)
von: Ugadiarov, Leonid, et al.
Veröffentlicht: (2026)
mini-vec2vec: Scaling Universal Geometry Alignment with Linear Transformations
von: Dar, Guy
Veröffentlicht: (2025)
von: Dar, Guy
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
von: Yang, Wenkai, et al.
Veröffentlicht: (2026)
A data-driven approach to modeling brain activity using differential equations
von: Andrey, Kuratov
Veröffentlicht: (2024)
von: Andrey, Kuratov
Veröffentlicht: (2024)
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling
von: Vejendla, Harshil
Veröffentlicht: (2025)
von: Vejendla, Harshil
Veröffentlicht: (2025)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
von: Havrilla, Alex, et al.
Veröffentlicht: (2024)
Large Language Models in the Task of Automatic Validation of Text Classifier Predictions
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Tsymbalov, Aleksandr, et al.
Veröffentlicht: (2025)
Catching rationalization in the act: detecting motivated reasoning before and after CoT via activation probing
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2026)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
Scaled and Inter-token Relation Enhanced Transformer for Sample-restricted Residential NILM
von: Rahman, Minhajur, et al.
Veröffentlicht: (2024)
von: Rahman, Minhajur, et al.
Veröffentlicht: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Associative Recurrent Memory Transformer
von: Rodkin, Ivan, et al.
Veröffentlicht: (2024) -
Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models
von: Chepurova, Alla, et al.
Veröffentlicht: (2025) -
In Search of Needles in a 11M Haystack: Recurrent Memory Finds What LLMs Miss
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024) -
Cramming 1568 Tokens into a Single Vector and Back Again: Exploring the Limits of Embedding Space Capacity
von: Kuratov, Yuri, et al.
Veröffentlicht: (2025) -
BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack
von: Kuratov, Yuri, et al.
Veröffentlicht: (2024)