Neuralink: Fast LLM Inference on Smartphones with Neuron Co-Activation Linking
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Tuowei, Fan, Ruwen, Huang, Minxing, Hao, Zixu, Li, Kun, Cao, Ting, Lu, Youyou, Zhang, Yaoxue, Ren, Ju |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
by: Wang, Tuowei, et al.
Published: (2024)
by: Wang, Tuowei, et al.
Published: (2024)
CORD: Co-design of Resource Allocation and Deadline Decomposition with Generative Profiling
by: Gifford, Robert, et al.
Published: (2025)
by: Gifford, Robert, et al.
Published: (2025)
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026)
by: Wang, Yirong, et al.
Published: (2026)
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
by: Jiang, Jinhu, et al.
Published: (2024)
by: Jiang, Jinhu, et al.
Published: (2024)
Delegation with Trust<T>: A Scalable, Type- and Memory-Safe Alternative to Locks
by: Ahmad, Noaman, et al.
Published: (2024)
by: Ahmad, Noaman, et al.
Published: (2024)
Columbo: Low Level End-to-End System Traces through Modular Full-System Simulation
by: Görgen, Jakob, et al.
Published: (2024)
by: Görgen, Jakob, et al.
Published: (2024)
Inspection of I/O Operations from System Call Traces using Directly-Follows-Graph
by: Sankaran, Aravind, et al.
Published: (2024)
by: Sankaran, Aravind, et al.
Published: (2024)
Optimizing System Memory Bandwidth with Micron CXL Memory Expansion Modules on Intel Xeon 6 Processors
by: Sehgal, Rohit, et al.
Published: (2024)
by: Sehgal, Rohit, et al.
Published: (2024)
Performance Characterization of AutoNUMA Memory Tiering on Graph Analytics
by: Moura, Diego, et al.
Published: (2022)
by: Moura, Diego, et al.
Published: (2022)
Data-Driven Power Modeling and Monitoring via Hardware Performance Counters Tracking
by: Mazzola, Sergio, et al.
Published: (2024)
by: Mazzola, Sergio, et al.
Published: (2024)
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
by: Hao, Zixu, et al.
Published: (2025)
by: Hao, Zixu, et al.
Published: (2025)
MNN-AECS: Energy Optimization for LLM Decoding on Mobile Devices via Adaptive Core Selection
by: Huang, Zhengxiang, et al.
Published: (2025)
by: Huang, Zhengxiang, et al.
Published: (2025)
SemaTune: Semantic-Aware Online OS Tuning with Large Language Models
by: Liargkovas, Georgios, et al.
Published: (2026)
by: Liargkovas, Georgios, et al.
Published: (2026)
CXLMemSim: A pure software simulated CXL.mem for performance characterization
by: Yang, Yiwei, et al.
Published: (2023)
by: Yang, Yiwei, et al.
Published: (2023)
Tidying Up the Address Space
by: Banakar, Vinay, et al.
Published: (2025)
by: Banakar, Vinay, et al.
Published: (2025)
Putting the Context back into Memory
by: Roberts, David A.
Published: (2025)
by: Roberts, David A.
Published: (2025)
A Limits Study of Memory-side Tiering Telemetry
by: Petrucci, Vinicius, et al.
Published: (2025)
by: Petrucci, Vinicius, et al.
Published: (2025)
CounterPoint: Using Hardware Event Counters to Refute and Refine Microarchitectural Assumptions (Extended Version)
by: Lindsay, Nick, et al.
Published: (2026)
by: Lindsay, Nick, et al.
Published: (2026)
Blink: CPU-Free LLM Inference by Delegating the Serving Stack to GPU and SmartNIC
by: Siavashi, Mohammad, et al.
Published: (2026)
by: Siavashi, Mohammad, et al.
Published: (2026)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
PipeANN-Filter: An Efficient Filtered Vector Search System on SSD
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Sensifi: A Wireless Sensing System for Ultra-High-Rate Applications
by: Li, Chia-Chi, et al.
Published: (2020)
by: Li, Chia-Chi, et al.
Published: (2020)
Characterizing Physical Memory Fragmentation
by: Mansi, Mark, et al.
Published: (2024)
by: Mansi, Mark, et al.
Published: (2024)
Energy-Aware CPU Orchestration in O-RAN: A dApp-Driven Lightweight Approach
by: Crespo, Francisco, et al.
Published: (2025)
by: Crespo, Francisco, et al.
Published: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
Mitigating GIL Bottlenecks in Edge AI Systems
by: Mandal, Mridankan, et al.
Published: (2026)
by: Mandal, Mridankan, et al.
Published: (2026)
ASC-Hook: fast and transparent system call hook for Arm
by: Shen, Yang, et al.
Published: (2024)
by: Shen, Yang, et al.
Published: (2024)
RAID Organizations for Improved Reliability and Performance: A Not Entirely Unbiased Tutorial (1st revision)
by: Thomasian, Alexander
Published: (2024)
by: Thomasian, Alexander
Published: (2024)
SwitchFS: Asynchronous Metadata Updates for Distributed Filesystems with In-Network Coordination
by: Xu, Jingwei, et al.
Published: (2024)
by: Xu, Jingwei, et al.
Published: (2024)
CPU-Limits kill Performance: Time to rethink Resource Control
by: Shetty, Chirag, et al.
Published: (2025)
by: Shetty, Chirag, et al.
Published: (2025)
A TRRIP Down Memory Lane: Temperature-Based Re-Reference Interval Prediction For Instruction Caching
by: Kao, Henry, et al.
Published: (2025)
by: Kao, Henry, et al.
Published: (2025)
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
by: Zhu, Yifan, et al.
Published: (2026)
by: Zhu, Yifan, et al.
Published: (2026)
Demystifying Serverless Costs on Public Platforms: Bridging Billing, Architecture, and OS Scheduling
by: Lin, Changyuan, et al.
Published: (2025)
by: Lin, Changyuan, et al.
Published: (2025)
Optimizing CPU Cache Utilization in Cloud VMs with Accurate Cache Abstraction
by: Tofigh, Mani, et al.
Published: (2025)
by: Tofigh, Mani, et al.
Published: (2025)
GPUVM: GPU-driven Unified Virtual Memory
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
by: Nazaraliyev, Nurlan, et al.
Published: (2024)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
The Impact of Private Equity and Venture Capital Funds on post-IPO Operational and Financial Performance in Brazilian invested companies
by: Bianca Piloto Sincerre
Published: (2019)
by: Bianca Piloto Sincerre
Published: (2019)
Understanding and Enhancing Linux Kernel-based Packet Switching on WiFi Access Points
by: Zhang, Shiqi, et al.
Published: (2024)
by: Zhang, Shiqi, et al.
Published: (2024)
Scheduling the Unschedulable: Taming Black-Box LLM Inference at Scale
by: Yuan, Renzhong, et al.
Published: (2026)
by: Yuan, Renzhong, et al.
Published: (2026)
Similar Items
-
Swarm: Co-Activation Aware KVCache Offloading Across Multiple SSDs
by: Wang, Tuowei, et al.
Published: (2026) -
Matryoshka: Optimization of Dynamic Diverse Quantum Chemistry Systems via Elastic Parallelism Transformation
by: Wang, Tuowei, et al.
Published: (2024) -
CORD: Co-design of Resource Allocation and Deadline Decomposition with Generative Profiling
by: Gifford, Robert, et al.
Published: (2025) -
2DIO: A Cache-Accurate Storage Microbenchmark
by: Wang, Yirong, et al.
Published: (2026) -
A System-Level Dynamic Binary Translator using Automatically-Learned Translation Rules
by: Jiang, Jinhu, et al.
Published: (2024)