Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
Fuente:
arXiv
Salvato in:
| Autori principali: | Jia, Fucheng, Wu, Zewen, Jiang, Shiqi, Jiang, Huiqiang, Zhang, Qianxi, Yang, Yuqing, Liu, Yunxin, Ren, Ju, Zhang, Deyu, Cao, Ting |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations
di: Jia, Fucheng, et al.
Pubblicazione: (2023)
di: Jia, Fucheng, et al.
Pubblicazione: (2023)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
di: Hao, Zixu, et al.
Pubblicazione: (2025)
di: Hao, Zixu, et al.
Pubblicazione: (2025)
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
di: Ding, Xin, et al.
Pubblicazione: (2025)
di: Ding, Xin, et al.
Pubblicazione: (2025)
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
di: Zheng, Yikai, et al.
Pubblicazione: (2026)
di: Zheng, Yikai, et al.
Pubblicazione: (2026)
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
di: Wu, Hao, et al.
Pubblicazione: (2024)
di: Wu, Hao, et al.
Pubblicazione: (2024)
KVNAND: Efficient On-Device Large Language Model Inference Using DRAM-Free In-Flash Computing
di: Deng, Lishuo, et al.
Pubblicazione: (2025)
di: Deng, Lishuo, et al.
Pubblicazione: (2025)
A First Look At Efficient And Secure On-Device LLM Inference Against KV Leakage
di: Yang, Huan, et al.
Pubblicazione: (2024)
di: Yang, Huan, et al.
Pubblicazione: (2024)
MemCompiler: Compile, Don't Inject -- State-Conditioned Memory for Embodied Agents
di: Ding, Xin, et al.
Pubblicazione: (2026)
di: Ding, Xin, et al.
Pubblicazione: (2026)
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
di: Li, Mo, et al.
Pubblicazione: (2025)
di: Li, Mo, et al.
Pubblicazione: (2025)
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
di: Ju, Ruofei, et al.
Pubblicazione: (2026)
Accelerating Prefilling via Decoding-time Contribution Sparsity
di: He, Zhiyuan, et al.
Pubblicazione: (2025)
di: He, Zhiyuan, et al.
Pubblicazione: (2025)
MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
di: Li, Wenxuan, et al.
Pubblicazione: (2025)
MobiFuse: A High-Precision On-device Depth Perception System with Multi-Data Fusion
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
di: Zhang, Jinrui, et al.
Pubblicazione: (2024)
SwapNet: Efficient Swapping for DNN Inference on Edge AI Devices Beyond the Memory Budget
di: Wang, Kun, et al.
Pubblicazione: (2024)
di: Wang, Kun, et al.
Pubblicazione: (2024)
LongLLMLingua: Accelerating and Enhancing LLMs in Long Context Scenarios via Prompt Compression
di: Jiang, Huiqiang, et al.
Pubblicazione: (2023)
di: Jiang, Huiqiang, et al.
Pubblicazione: (2023)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
SortedRL: Accelerating RL Training for LLMs through Online Length-Aware Scheduling
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
di: Zhang, Yiqi, et al.
Pubblicazione: (2026)
FP-Rowhammer: DRAM-Based Device Fingerprinting
di: Venugopalan, Hari, et al.
Pubblicazione: (2023)
di: Venugopalan, Hari, et al.
Pubblicazione: (2023)
JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
di: Cai, Aichen, et al.
Pubblicazione: (2026)
di: Cai, Aichen, et al.
Pubblicazione: (2026)
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding
di: Zhang, Yike, et al.
Pubblicazione: (2025)
di: Zhang, Yike, et al.
Pubblicazione: (2025)
Revisiting DRAM Read Disturbance: Identifying Inconsistencies Between Experimental Characterization and Device-Level Studies
di: Luo, Haocong, et al.
Pubblicazione: (2025)
di: Luo, Haocong, et al.
Pubblicazione: (2025)
Settling Weighted Token Swapping up to Algorithmic Barriers
di: Wein, Nicole, et al.
Pubblicazione: (2025)
di: Wein, Nicole, et al.
Pubblicazione: (2025)
Weisfeiler‐Lehman Kernel Augmented Product Representation for Queries on Large‐Scale BIM Scenes
di: Huiqiang Hu, et al.
Pubblicazione: (2025)
di: Huiqiang Hu, et al.
Pubblicazione: (2025)
FlashRecovery: Fast and Low-Cost Recovery from Failures for Large-Scale Training of LLMs
di: Zhang, Haijun, et al.
Pubblicazione: (2025)
di: Zhang, Haijun, et al.
Pubblicazione: (2025)
Distinguishing Translations by Human, NMT, and ChatGPT: A Linguistic and Statistical Approach
di: Jiang, Zhaokun, et al.
Pubblicazione: (2023)
di: Jiang, Zhaokun, et al.
Pubblicazione: (2023)
Convergences and Divergences between Automatic Assessment and Human Evaluation: Insights from Comparing ChatGPT-Generated Translation and Neural Machine Translation
di: Jiang, Zhaokun, et al.
Pubblicazione: (2024)
di: Jiang, Zhaokun, et al.
Pubblicazione: (2024)
RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
di: Liu, Di, et al.
Pubblicazione: (2024)
di: Liu, Di, et al.
Pubblicazione: (2024)
Hardware Co-Design Scaling Laws via Roofline Modelling for On-Device LLMs
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
di: Sun, Luoyang, et al.
Pubblicazione: (2026)
Uncovering the Role of Initial Saliency in U-Shaped Attention Bias: Scaling Initial Token Weight for Enhanced Long-Text Processing
di: Qiang, Zewen, et al.
Pubblicazione: (2025)
di: Qiang, Zewen, et al.
Pubblicazione: (2025)
Position Engineering: Boosting Large Language Models through Positional Information Manipulation
di: He, Zhiyuan, et al.
Pubblicazione: (2024)
di: He, Zhiyuan, et al.
Pubblicazione: (2024)
An Analytic Model to Determine the Interstitial-Solute Energetics and Underlying Mechanism in Refractory High-Entropy Alloys
di: Zhu, Qianxi, et al.
Pubblicazione: (2025)
di: Zhu, Qianxi, et al.
Pubblicazione: (2025)
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
di: Li, Xiangyu, et al.
Pubblicazione: (2025)
Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
di: Qian, Jiaxu, et al.
Pubblicazione: (2025)
di: Qian, Jiaxu, et al.
Pubblicazione: (2025)
ProRe: A Proactive Reward System for GUI Agents via Reasoner-Actor Collaboration
di: Dai, Gaole, et al.
Pubblicazione: (2025)
di: Dai, Gaole, et al.
Pubblicazione: (2025)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment
di: Dai, Gaole, et al.
Pubblicazione: (2025)
di: Dai, Gaole, et al.
Pubblicazione: (2025)
AVA: Towards Agentic Video Analytics with Vision Language Models
di: Yan, Yuxuan, et al.
Pubblicazione: (2025)
di: Yan, Yuxuan, et al.
Pubblicazione: (2025)
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
di: Li, Yucheng, et al.
Pubblicazione: (2025)
di: Li, Yucheng, et al.
Pubblicazione: (2025)
Mitigate Position Bias in Large Language Models via Scaling a Single Dimension
di: Yu, Yijiong, et al.
Pubblicazione: (2024)
di: Yu, Yijiong, et al.
Pubblicazione: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Empowering In-Browser Deep Learning Inference on Edge Devices with Just-in-Time Kernel Optimizations
di: Jia, Fucheng, et al.
Pubblicazione: (2023) -
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
di: Hao, Zixu, et al.
Pubblicazione: (2025) -
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
di: Ding, Xin, et al.
Pubblicazione: (2025) -
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
di: Zheng, Yikai, et al.
Pubblicazione: (2026) -
Making Every Frame Matter: Continuous Activity Recognition in Streaming Video via Adaptive Video Context Modeling
di: Wu, Hao, et al.
Pubblicazione: (2024)