LLMPi: Optimizing LLMs for High-Throughput on Raspberry Pi
Fuente:
arXiv
Saved in:
| Main Authors: | Ardakani, Mahsa, Malekar, Jinendra, Zand, Ramtin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Matmul or No Matmul in the Era of 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2024)
by: Malekar, Jinendra, et al.
Published: (2024)
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025)
by: Malekar, Jinendra, et al.
Published: (2025)
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024)
by: Seekings, James, et al.
Published: (2024)
High-Throughput SAT Sampling
by: Ardakani, Arash, et al.
Published: (2025)
by: Ardakani, Arash, et al.
Published: (2025)
Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural Networks
by: King, Grant, et al.
Published: (2025)
by: King, Grant, et al.
Published: (2025)
Flex-TPU: A Flexible TPU with Runtime Reconfigurable Dataflow Architecture
by: Elbtity, Mohammed, et al.
Published: (2024)
by: Elbtity, Mohammed, et al.
Published: (2024)
NSF-MAP: Neurosymbolic Multimodal Fusion for Robust and Interpretable Anomaly Prediction in Assembly Pipelines
by: Shyalika, Chathurangi, et al.
Published: (2025)
by: Shyalika, Chathurangi, et al.
Published: (2025)
PiCO: Peer Review in LLMs based on the Consistency Optimization
by: Ning, Kun-Peng, et al.
Published: (2024)
by: Ning, Kun-Peng, et al.
Published: (2024)
Self Paced Gaussian Contextual Reinforcement Learning
by: Ardakani, Mohsen Sahraei, et al.
Published: (2026)
by: Ardakani, Mohsen Sahraei, et al.
Published: (2026)
Inverse Reinforcement Learning by Estimating Expertise of Demonstrators
by: Beliaev, Mark, et al.
Published: (2024)
by: Beliaev, Mark, et al.
Published: (2024)
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
HD-PiSSA: High-Rank Distributed Orthogonal Adaptation
by: Wang, Yiding, et al.
Published: (2025)
by: Wang, Yiding, et al.
Published: (2025)
Scalable Option Learning in High-Throughput Environments
by: Henaff, Mikael, et al.
Published: (2025)
by: Henaff, Mikael, et al.
Published: (2025)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024)
by: Mohammadi, Seyedali, et al.
Published: (2024)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Conflict-Aware Adversarial Training
by: Xue, Zhiyu, et al.
Published: (2024)
by: Xue, Zhiyu, et al.
Published: (2024)
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
by: Shrestha, Susav, et al.
Published: (2025)
by: Shrestha, Susav, et al.
Published: (2025)
PiCSRL: Physics-Informed Contextual Spectral Reinforcement Learning
by: Azadani, Mitra Nasr, et al.
Published: (2026)
by: Azadani, Mitra Nasr, et al.
Published: (2026)
PiFlow: Principle-Aware Scientific Discovery with Multi-Agent Collaboration
by: Pu, Yingming, et al.
Published: (2025)
by: Pu, Yingming, et al.
Published: (2025)
PiCa: Parameter-Efficient Fine-Tuning with Column Space Projection
by: Hwang, Junseo, et al.
Published: (2025)
by: Hwang, Junseo, et al.
Published: (2025)
Level Generation with Constrained Expressive Range
by: Bazzaz, Mahsa, et al.
Published: (2025)
by: Bazzaz, Mahsa, et al.
Published: (2025)
Analysis of Robustness of a Large Game Corpus
by: Bazzaz, Mahsa, et al.
Published: (2025)
by: Bazzaz, Mahsa, et al.
Published: (2025)
Guided Game Level Repair via Explainable AI
by: Bazzaz, Mahsa, et al.
Published: (2024)
by: Bazzaz, Mahsa, et al.
Published: (2024)
Controllable Game Level Generation: Assessing the Effect of Negative Examples in GAN Models
by: Bazzaz, Mahsa, et al.
Published: (2024)
by: Bazzaz, Mahsa, et al.
Published: (2024)
EnviroPiNet: A Physics-Guided AI Model for Predicting Biofilter Performance
by: Uzma, et al.
Published: (2025)
by: Uzma, et al.
Published: (2025)
PiSSA: Principal Singular Values and Singular Vectors Adaptation of Large Language Models
by: Meng, Fanxu, et al.
Published: (2024)
by: Meng, Fanxu, et al.
Published: (2024)
PiXTime: A Model for Federated Time Series Forecasting with Heterogeneous Data across Nodes
by: Zhou, Yiming, et al.
Published: (2026)
by: Zhou, Yiming, et al.
Published: (2026)
PiDR: Physics-Informed Inertial Dead Reckoning for Autonomous Platforms
by: Sahoo, Arup Kumar, et al.
Published: (2026)
by: Sahoo, Arup Kumar, et al.
Published: (2026)
Enabling High Data Throughput Reinforcement Learning on GPUs: A Domain Agnostic Framework for Data-Driven Scientific Research
by: Lan, Tian, et al.
Published: (2024)
by: Lan, Tian, et al.
Published: (2024)
Implementation of Google Assistant & Amazon Alexa on Raspberry Pi
by: Arya, Shailesh D., et al.
Published: (2020)
by: Arya, Shailesh D., et al.
Published: (2020)
Understanding the Challenges in Iterative Generative Optimization with LLMs
by: Nie, Allen, et al.
Published: (2026)
by: Nie, Allen, et al.
Published: (2026)
Aligning CodeLLMs with Direct Preference Optimization
by: Miao, Yibo, et al.
Published: (2024)
by: Miao, Yibo, et al.
Published: (2024)
Calibration-Aware Policy Optimization for Reasoning LLMs
by: Wang, Ziqi, et al.
Published: (2026)
by: Wang, Ziqi, et al.
Published: (2026)
BI-DCGAN: A Theoretically Grounded Bayesian Framework for Efficient and Diverse GANs
by: Valizadeh, Mahsa, et al.
Published: (2025)
by: Valizadeh, Mahsa, et al.
Published: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Litespark Technical Report: High-Throughput, Energy-Efficient LLM Training Framework
by: Dade, Nii Osae Osae, et al.
Published: (2025)
by: Dade, Nii Osae Osae, et al.
Published: (2025)
From Theory to Throughput: CUDA-Optimized APML for Large-Batch 3D Learning
by: Sharifipour, Sasan, et al.
Published: (2025)
by: Sharifipour, Sasan, et al.
Published: (2025)
Accelerating LLM Inference Throughput via Asynchronous KV Cache Prefetching
by: Dong, Yanhao, et al.
Published: (2025)
by: Dong, Yanhao, et al.
Published: (2025)
AI-assisted Protocol Information Extraction For Improved Accuracy and Efficiency in Clinical Trial Workflows
by: Babaeipour, Ramtin, et al.
Published: (2026)
by: Babaeipour, Ramtin, et al.
Published: (2026)
Spiffy: Efficient Implementation of CoLaNET for Raspberry Pi
by: Derzhavin, Andrey, et al.
Published: (2025)
by: Derzhavin, Andrey, et al.
Published: (2025)
Similar Items
-
Matmul or No Matmul in the Era of 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2024) -
PIM-LLM: A High-Throughput Hybrid PIM Architecture for 1-bit LLMs
by: Malekar, Jinendra, et al.
Published: (2025) -
Towards Efficient Deployment of Hybrid SNNs on Neuromorphic and Edge AI Hardware
by: Seekings, James, et al.
Published: (2024) -
High-Throughput SAT Sampling
by: Ardakani, Arash, et al.
Published: (2025) -
Rep Smarter, Not Harder: AI Hypertrophy Coaching with Wearable Sensors and Edge Neural Networks
by: King, Grant, et al.
Published: (2025)