Building on Efficient Foundations: Effectively Training LLMs with Structured Feedforward Layers
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Xiuying, Moalla, Skander, Pascanu, Razvan, Gulcehre, Caglar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024)
by: Wei, Xiuying, et al.
Published: (2024)
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025)
by: Wei, Xiuying, et al.
Published: (2025)
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
by: Moalla, Skander, et al.
Published: (2024)
by: Moalla, Skander, et al.
Published: (2024)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
by: Matrenok, Simon, et al.
Published: (2025)
by: Matrenok, Simon, et al.
Published: (2025)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
Promises, Outlooks and Challenges of Diffusion Language Modeling
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
by: Wei, Xiuying, et al.
Published: (2026)
by: Wei, Xiuying, et al.
Published: (2026)
Python Machine Learning Research Template
by: Moalla, Skander
Published: (2025)
by: Moalla, Skander
Published: (2025)
Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues
by: Orvieto, Antonio, et al.
Published: (2023)
by: Orvieto, Antonio, et al.
Published: (2023)
The Illusion of Stochasticity in LLMs
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Why do LLMs attend to the first token?
by: Barbero, Federico, et al.
Published: (2025)
by: Barbero, Federico, et al.
Published: (2025)
Aligning Large Language Models with Diverse Political Viewpoints
by: Stammbach, Dominik, et al.
Published: (2024)
by: Stammbach, Dominik, et al.
Published: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata
by: Schurger-Foy, Adrien, et al.
Published: (2025)
by: Schurger-Foy, Adrien, et al.
Published: (2025)
Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models
by: Gu, Xiangming, et al.
Published: (2026)
by: Gu, Xiangming, et al.
Published: (2026)
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
by: De, Soham, et al.
Published: (2024)
by: De, Soham, et al.
Published: (2024)
Round and Round We Go! What makes Rotary Positional Encodings useful?
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Exploring the Ethical Concerns in User Reviews of Mental Health Apps using Topic Modeling and Sentiment Analysis
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
Perplexity Cannot Always Tell Right from Wrong
by: Veličković, Petar, et al.
Published: (2026)
by: Veličković, Petar, et al.
Published: (2026)
Fleet of Agents: Coordinated Problem Solving with Large Language Models
by: Klein, Lars, et al.
Published: (2024)
by: Klein, Lars, et al.
Published: (2024)
How do language models learn facts? Dynamics, curricula and hallucinations
by: Zucchet, Nicolas, et al.
Published: (2025)
by: Zucchet, Nicolas, et al.
Published: (2025)
Self-Recognition in Language Models
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
FreeCtrl: Constructing Control Centers with Feedforward Layers for Learning-Free Controllable Text Generation
by: Feng, Zijian, et al.
Published: (2024)
by: Feng, Zijian, et al.
Published: (2024)
SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
by: Xie, Chengxing, et al.
Published: (2025)
by: Xie, Chengxing, et al.
Published: (2025)
Layer-Wise Quantization: A Pragmatic and Effective Method for Quantizing LLMs Beyond Integer Bit-Levels
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
LayerBoost: Layer-Aware Attention Reduction for Efficient LLMs
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
by: Souibgui, Mohamed Ali, et al.
Published: (2026)
Building Efficient and Effective OpenQA Systems for Low-Resource Languages
by: Budur, Emrah, et al.
Published: (2024)
by: Budur, Emrah, et al.
Published: (2024)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
by: Wang, Chenxi, et al.
Published: (2025)
by: Wang, Chenxi, et al.
Published: (2025)
Flipping Knowledge Distillation: Leveraging Small Models' Expertise to Enhance LLMs in Text Matching
by: Li, Mingzhe, et al.
Published: (2025)
by: Li, Mingzhe, et al.
Published: (2025)
Transformers meet Neural Algorithmic Reasoners
by: Bounsi, Wilfried, et al.
Published: (2024)
by: Bounsi, Wilfried, et al.
Published: (2024)
GRASP LoRA: GRPO Guided Adapter Sparsity Policy for Cross Lingual Transfer
by: Hassan, Besher, et al.
Published: (2026)
by: Hassan, Besher, et al.
Published: (2026)
GeneMamba: An Efficient and Effective Foundation Model on Single Cell Data
by: Qi, Cong, et al.
Published: (2025)
by: Qi, Cong, et al.
Published: (2025)
Transformers need glasses! Information over-squashing in language tasks
by: Barbero, Federico, et al.
Published: (2024)
by: Barbero, Federico, et al.
Published: (2024)
Configurable Foundation Models: Building LLMs from a Modular Perspective
by: Xiao, Chaojun, et al.
Published: (2024)
by: Xiao, Chaojun, et al.
Published: (2024)
ELO: Efficient Layer-Specific Optimization for Continual Pretraining of Multilingual LLMs
by: Yoo, HanGyeol, et al.
Published: (2026)
by: Yoo, HanGyeol, et al.
Published: (2026)
Beyond Profile: From Surface-Level Facts to Deep Persona Simulation in LLMs
by: Wang, Zixiao, et al.
Published: (2025)
by: Wang, Zixiao, et al.
Published: (2025)
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
by: Feng, Ruixiang, et al.
Published: (2025)
by: Feng, Ruixiang, et al.
Published: (2025)
Towards Building Efficient Sentence BERT Models using Layer Pruning
by: Shelke, Anushka, et al.
Published: (2024)
by: Shelke, Anushka, et al.
Published: (2024)
GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
Similar Items
-
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis
by: Wei, Xiuying, et al.
Published: (2024) -
RAT: Bridging RNN Efficiency and Attention Accuracy via Chunk-based Sequence Modeling
by: Wei, Xiuying, et al.
Published: (2025) -
No Representation, No Trust: Connecting Representation, Collapse, and Trust Issues in PPO
by: Moalla, Skander, et al.
Published: (2024) -
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
by: Matrenok, Simon, et al.
Published: (2025) -
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
by: Deschenaux, Justin, et al.
Published: (2024)