LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Fuente:
arXiv
Salvato in:
| Autori principali: | Alizadeh, Keivan, Mirzadeh, Iman, Belenko, Dmitry, Khatamifard, Karen, Cho, Minsik, Del Mundo, Carlo C, Rastegari, Mohammad, Farajtabar, Mehrdad |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
di: Alizadeh, Keivan, et al.
Pubblicazione: (2024)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2024)
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
di: Mirzadeh, Iman, et al.
Pubblicazione: (2024)
di: Mirzadeh, Iman, et al.
Pubblicazione: (2024)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
di: Samragh, Mohammad, et al.
Pubblicazione: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
di: Shojaee, Parshin, et al.
Pubblicazione: (2025)
di: Shojaee, Parshin, et al.
Pubblicazione: (2025)
Computational Bottlenecks of Training Small-scale Large Language Models
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
di: Ashkboos, Saleh, et al.
Pubblicazione: (2024)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
di: Cho, Minsik, et al.
Pubblicazione: (2024)
di: Cho, Minsik, et al.
Pubblicazione: (2024)
OpenELM: An Efficient Language Model Family with Open Training and Inference Framework
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
Barriers for Learning in an Evolving World: Mathematical Understanding of Loss of Plasticity
di: Joudaki, Amir, et al.
Pubblicazione: (2025)
di: Joudaki, Amir, et al.
Pubblicazione: (2025)
SALSA: Soup-based Alignment Learning for Stronger Adaptation in RLHF
di: Chegini, Atoosa, et al.
Pubblicazione: (2024)
di: Chegini, Atoosa, et al.
Pubblicazione: (2024)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
di: Fu, Qichen, et al.
Pubblicazione: (2024)
di: Fu, Qichen, et al.
Pubblicazione: (2024)
MemoryLLM: Plug-n-Play Interpretable Feed-Forward Memory for Transformers
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
di: Samragh, Mohammad, et al.
Pubblicazione: (2025)
Knowledge Transfer from Vision Foundation Models for Efficient Training of Small Task-specific Models
di: Vemulapalli, Raviteja, et al.
Pubblicazione: (2023)
di: Vemulapalli, Raviteja, et al.
Pubblicazione: (2023)
TIDE: Every Layer Knows the Token Beneath the Context
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
di: Jaiswal, Ajay, et al.
Pubblicazione: (2026)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
di: Hannah, Lauren. A, et al.
Pubblicazione: (2025)
TiC-LM: A Web-Scale Benchmark for Time-Continual LLM Pretraining
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
di: Li, Jeffrey, et al.
Pubblicazione: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
di: Hoang, Duc N. M, et al.
Pubblicazione: (2023)
CatLIP: CLIP-level Visual Recognition Accuracy with 2.7x Faster Pre-training on Web-scale Image-Text Data
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
di: Mehta, Sachin, et al.
Pubblicazione: (2024)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
di: Kim, Han-Byul, et al.
Pubblicazione: (2025)
Towards Low-bit Communication for Tensor Parallel LLM Inference
di: Dong, Harry, et al.
Pubblicazione: (2024)
di: Dong, Harry, et al.
Pubblicazione: (2024)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
di: Nishu, Kumari, et al.
Pubblicazione: (2025)
di: Nishu, Kumari, et al.
Pubblicazione: (2025)
LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2026)
di: Mozaffari, Mohammad, et al.
Pubblicazione: (2026)
SAM-CLIP: Merging Vision Foundation Models towards Semantic and Spatial Understanding
di: Wang, Haoxiang, et al.
Pubblicazione: (2023)
di: Wang, Haoxiang, et al.
Pubblicazione: (2023)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
di: Armandpour, Mohammadreza, et al.
Pubblicazione: (2026)
DeFi Liquidation Risk Modeling Using Geometric Brownian Motion
di: Belenko, Timofei, et al.
Pubblicazione: (2025)
di: Belenko, Timofei, et al.
Pubblicazione: (2025)
Speculative Streaming: Fast LLM Inference without Auxiliary Models
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2024)
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2024)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
The path towards contact-based physical human-robot interaction
di: Farajtabar, Mohammad, et al.
Pubblicazione: (2024)
di: Farajtabar, Mohammad, et al.
Pubblicazione: (2024)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
di: Zibakhsh, Soheil, et al.
Pubblicazione: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
di: Hoang, Duc, et al.
Pubblicazione: (2026)
di: Hoang, Duc, et al.
Pubblicazione: (2026)
LUT-LLM: Efficient Large Language Model Inference with Memory-based Computations on FPGAs
di: He, Zifan, et al.
Pubblicazione: (2025)
di: He, Zifan, et al.
Pubblicazione: (2025)
SpecHop: Continuous Speculation for Accelerating Multi-Hop Retrieval Agents
di: Saberi, Mehrdad, et al.
Pubblicazione: (2026)
di: Saberi, Mehrdad, et al.
Pubblicazione: (2026)
The hybrid dilemma -- do hybrid technologies play a transitionary or stationary role in transitions processes?
di: Phirouzabadi, Amir Mirzadeh
Pubblicazione: (2025)
di: Phirouzabadi, Amir Mirzadeh
Pubblicazione: (2025)
SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
di: Nishu, Kumari, et al.
Pubblicazione: (2024)
di: Nishu, Kumari, et al.
Pubblicazione: (2024)
Test Time Training for AC Power Flow Surrogates via Physics and Operational Constraint Refinement
di: Dogoulis, Panteleimon, et al.
Pubblicazione: (2025)
di: Dogoulis, Panteleimon, et al.
Pubblicazione: (2025)
Achieving Fine‐Grained Microstructure in Low‐Alloy Steel: A Study on Static Recrystallization Using Experimental and Simulation Approaches
di: Mahdiyeh Baharvand, et al.
Pubblicazione: (2025)
di: Mahdiyeh Baharvand, et al.
Pubblicazione: (2025)
Low-rank Momentum Factorization for Memory Efficient Training
di: Mahdavinia, Pouria, et al.
Pubblicazione: (2025)
di: Mahdavinia, Pouria, et al.
Pubblicazione: (2025)
Resource-Efficient Iterative LLM-Based NAS with Feedback Memory
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
di: Gu, Xiaojie, et al.
Pubblicazione: (2026)
Harnessing the Ecological and Genomic Adaptability of the Bacterial Genus Massilia for Environmental and Industrial Applications
di: Kamyar Amirhosseini, et al.
Pubblicazione: (2025)
di: Kamyar Amirhosseini, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
di: Alizadeh, Keivan, et al.
Pubblicazione: (2024) -
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
di: Mirzadeh, Iman, et al.
Pubblicazione: (2024) -
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
di: Samragh, Mohammad, et al.
Pubblicazione: (2024) -
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
di: Alizadeh, Keivan, et al.
Pubblicazione: (2026) -
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
di: Shojaee, Parshin, et al.
Pubblicazione: (2025)