Mixture of Chapters: Scaling Learnt Memory in Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tibrewal, Tasmay Pankaj, Saha, Pritish, Meda, Ankit, Singh, Kunal, Moturi, Pradeep |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
von: Singh, Kunal, et al.
Veröffentlicht: (2025)
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
von: Singh, Shreyas, et al.
Veröffentlicht: (2025)
von: Singh, Shreyas, et al.
Veröffentlicht: (2025)
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
von: Chen, Yilong, et al.
Veröffentlicht: (2026)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
von: Ildiz, M. Emrullah, et al.
Veröffentlicht: (2024)
MoM: Linear Sequence Modeling with Mixture-of-Memories
von: Du, Jusen, et al.
Veröffentlicht: (2025)
von: Du, Jusen, et al.
Veröffentlicht: (2025)
On the Spatial Structure of Mixture-of-Experts in Transformers
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
von: Bershatsky, Daniel, et al.
Veröffentlicht: (2025)
Scaling Laws for Fine-Grained Mixture of Experts
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
von: Krajewski, Jakub, et al.
Veröffentlicht: (2024)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
von: Basu, Soumya, et al.
Veröffentlicht: (2024)
An Evolved Universal Transformer Memory
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
von: Cetin, Edoardo, et al.
Veröffentlicht: (2024)
MobileMoE: Scaling On-Device Mixture of Experts
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
von: Chen, Yanbei, et al.
Veröffentlicht: (2026)
Towards a Comprehensive Scaling Law of Mixture-of-Experts
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
von: Zhao, Guoliang, et al.
Veröffentlicht: (2025)
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training
von: Mistry, Deven Mahesh, et al.
Veröffentlicht: (2025)
von: Mistry, Deven Mahesh, et al.
Veröffentlicht: (2025)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Clustering-driven Memory Compression for On-device Large Language Models
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
von: Bohdal, Ondrej, et al.
Veröffentlicht: (2026)
Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Models
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
von: Vendrell, Victor Conchello, et al.
Veröffentlicht: (2026)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2024)
von: Bandarkar, Lucas, et al.
Veröffentlicht: (2024)
Think Before You Act: Decision Transformers with Working Memory
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2026)
von: NVIDIA, et al.
Veröffentlicht: (2026)
MoVE: Mixture of Value Embeddings -- A New Axis for Scaling Parametric Memory in Autoregressive Models
von: Li, Yangyan
Veröffentlicht: (2026)
von: Li, Yangyan
Veröffentlicht: (2026)
Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
Working Memory Constraints Scaffold Learning in Transformers under Data Scarcity
von: Madhyastha, Pranava, et al.
Veröffentlicht: (2026)
von: Madhyastha, Pranava, et al.
Veröffentlicht: (2026)
SPINAL -- Scaling-law and Preference Integration in Neural Alignment Layers
von: Das, Arion, et al.
Veröffentlicht: (2026)
von: Das, Arion, et al.
Veröffentlicht: (2026)
MeKi: Memory-based Expert Knowledge Injection for Efficient LLM Scaling
von: Ding, Ning, et al.
Veröffentlicht: (2026)
von: Ding, Ning, et al.
Veröffentlicht: (2026)
Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
von: Ling Team, et al.
Veröffentlicht: (2025)
von: Ling Team, et al.
Veröffentlicht: (2025)
Unifying Learning Dynamics and Generalization in Transformers Scaling Law
von: Yang, Chiwun
Veröffentlicht: (2025)
von: Yang, Chiwun
Veröffentlicht: (2025)
Scaling Transformer to 1M tokens and beyond with RMT
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
von: Bulatov, Aydar, et al.
Veröffentlicht: (2023)
RadLite: Multi-Task LoRA Fine-Tuning of Small Language Models for CPU-Deployable Radiology AI
von: Gupta, Pankaj, et al.
Veröffentlicht: (2026)
von: Gupta, Pankaj, et al.
Veröffentlicht: (2026)
Remember This Event That Year? Assessing Temporal Information and Reasoning in Large Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
QuIM-RAG: Advancing Retrieval-Augmented Generation with Inverted Question Matching for Enhanced QA Performance
von: Saha, Binita, et al.
Veröffentlicht: (2025)
von: Saha, Binita, et al.
Veröffentlicht: (2025)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
von: Li, Yingcong, et al.
Veröffentlicht: (2024)
Probing the Limits of Compressive Memory: A Study of Infini-Attention in Small-Scale Pretraining
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2025)
Mixtures of In-Context Learners
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
von: Hong, Giwon, et al.
Veröffentlicht: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
von: Kamigaito, Hidetaka, et al.
Veröffentlicht: (2025)
Think before you speak: Training Language Models With Pause Tokens
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
von: Goyal, Sachin, et al.
Veröffentlicht: (2023)
Language Model Cascades: Token-level uncertainty and beyond
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
von: Gupta, Neha, et al.
Veröffentlicht: (2024)
Parameter-Efficient Routed Fine-Tuning: Mixture-of-Experts Demands Mixture of Adaptation Modules
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
von: Liu, Yilun, et al.
Veröffentlicht: (2025)
OLMoE: Open Mixture-of-Experts Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance
von: Singh, Kunal, et al.
Veröffentlicht: (2025) -
Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs
von: Singh, Shreyas, et al.
Veröffentlicht: (2025) -
Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
von: Ludziejewski, Jan, et al.
Veröffentlicht: (2025) -
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
von: Chen, Yilong, et al.
Veröffentlicht: (2026) -
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)