Gespeichert in:
| Hauptverfasser: | Jaiswal, Ajay, Hannah, Lauren, Kim, Han-Byul, Hoang, Duc, Kundu, Arnav, Farajtabar, Mehrdad, Cho, Minsik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.00398 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TIDE: Every Layer Knows the Token Beneath the Context
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2026)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2026)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
von: Kim, Minsoo, et al.
Veröffentlicht: (2025)
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025)
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
von: Samragh, Mohammad, et al.
Veröffentlicht: (2025)
von: Samragh, Mohammad, et al.
Veröffentlicht: (2025)
SpecMD: A Comprehensive Study On Speculative Expert Prefetching
von: Hoang, Duc, et al.
Veröffentlicht: (2026)
von: Hoang, Duc, et al.
Veröffentlicht: (2026)
Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
von: Armandpour, Mohammadreza, et al.
Veröffentlicht: (2026)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
M+: Extending MemoryLLM with Scalable Long-Term Memory
von: Wang, Yu, et al.
Veröffentlicht: (2025)
von: Wang, Yu, et al.
Veröffentlicht: (2025)
LLM in a flash: Efficient Large Language Model Inference with Limited Memory
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2023)
MoEs Are Stronger than You Think: Hyper-Parallel Inference Scaling with RoE
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
von: Zibakhsh, Soheil, et al.
Veröffentlicht: (2025)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
von: Bhendawade, Nikhil, et al.
Veröffentlicht: (2025)
R2 Loss: Range Restriction Loss for Model Compression and Quantization
von: Kundu, Arnav, et al.
Veröffentlicht: (2023)
von: Kundu, Arnav, et al.
Veröffentlicht: (2023)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2024)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
von: Samragh, Mohammad, et al.
Veröffentlicht: (2024)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
From Dense to Dynamic: Token-Difficulty Driven MoEfication of Pre-Trained LLMs
von: Nishu, Kumari, et al.
Veröffentlicht: (2025)
von: Nishu, Kumari, et al.
Veröffentlicht: (2025)
Do Compressed LLMs Forget Knowledge? An Experimental Study with Practical Implications
von: Hoang, Duc N. M, et al.
Veröffentlicht: (2023)
von: Hoang, Duc N. M, et al.
Veröffentlicht: (2023)
Leveraging Data to Say No: Memory Augmented Plug-and-Play Selective Prediction
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aditya, et al.
Veröffentlicht: (2026)
NGM: A Plug-and-Play Training-Free Memory Module for LLMs
von: Qu, Yuwen, et al.
Veröffentlicht: (2026)
von: Qu, Yuwen, et al.
Veröffentlicht: (2026)
Self-supervised Deep Hyperspectral Inpainting with the Plug and Play and Deep Image Prior Models
von: Li, Shuo, et al.
Veröffentlicht: (2025)
von: Li, Shuo, et al.
Veröffentlicht: (2025)
Analysis and Synthesis Denoisers for Forward-Backward Plug-and-Play Algorithms
von: Kowalski, Matthieu, et al.
Veröffentlicht: (2024)
von: Kowalski, Matthieu, et al.
Veröffentlicht: (2024)
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
von: Sarawgi, Utkarsh Oggy, et al.
Veröffentlicht: (2023)
von: Sarawgi, Utkarsh Oggy, et al.
Veröffentlicht: (2023)
Romanization-Induced Mispronunciations in Korean: How Latin Letters Alter the Perception of Japanese Voiceless Consonants
von: Kang, Byul
Veröffentlicht: (2025)
von: Kang, Byul
Veröffentlicht: (2025)
KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
von: Cho, Minsik, et al.
Veröffentlicht: (2024)
PlugMem: A Task-Agnostic Plugin Memory Module for LLM Agents
von: Yang, Ke, et al.
Veröffentlicht: (2026)
von: Yang, Ke, et al.
Veröffentlicht: (2026)
Uniform boundedness on rational maps with automorphisms
von: Han, Minsik
Veröffentlicht: (2024)
von: Han, Minsik
Veröffentlicht: (2024)
A Study of Student Dependency on Artificial Intelligence Applications in their Education: With Reference to Indore City
von: Ajay Jaiswal
Veröffentlicht: (2025)
von: Ajay Jaiswal
Veröffentlicht: (2025)
Topological transition as a percolation of the Berry curvature
von: Kim, Han-Byul, et al.
Veröffentlicht: (2024)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2024)
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
von: Kim, HyunJin, et al.
Veröffentlicht: (2023)
von: Kim, HyunJin, et al.
Veröffentlicht: (2023)
Online Temporal Action Localization with Memory-Augmented Transformer
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
MemOrb: A Plug-and-Play Verbal-Reinforcement Memory Layer for E-Commerce Customer Service
von: Huang, Yizhe, et al.
Veröffentlicht: (2025)
von: Huang, Yizhe, et al.
Veröffentlicht: (2025)
Towards Low-bit Communication for Tensor Parallel LLM Inference
von: Dong, Harry, et al.
Veröffentlicht: (2024)
von: Dong, Harry, et al.
Veröffentlicht: (2024)
Safe Memory Reclamation Techniques
von: Singh, Ajay
Veröffentlicht: (2025)
von: Singh, Ajay
Veröffentlicht: (2025)
F4Splat: Feed-Forward Predictive Densification for Feed-Forward 3D Gaussian Splatting
von: Kim, Injae, et al.
Veröffentlicht: (2026)
von: Kim, Injae, et al.
Veröffentlicht: (2026)
ProTransformer: Robustify Transformers via Plug-and-Play Paradigm
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
von: Hou, Zhichao, et al.
Veröffentlicht: (2024)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
Plug-and-Play Transformer Modules for Test-Time Adaptation
von: Chang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
TIDE: Every Layer Knows the Token Beneath the Context
von: Jaiswal, Ajay, et al.
Veröffentlicht: (2026) -
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025) -
EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained Environments
von: Kim, Minsoo, et al.
Veröffentlicht: (2025) -
MoE-PHDS: One MoE checkpoint for flexible runtime sparsity
von: Hannah, Lauren. A, et al.
Veröffentlicht: (2025) -
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
von: Samragh, Mohammad, et al.
Veröffentlicht: (2025)