An Evaluation of LLMs Inference on Popular Single-board Computers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tung, Nguyen, Nguyen, Tuyen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
von: Song, Jingwei, et al.
Veröffentlicht: (2025)
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
von: Liu, Wentao, et al.
Veröffentlicht: (2025)
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025)
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026)
LLM as HPC Expert: Extending RAG Architecture for HPC Data
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)
Benchmarking Federated Learning in Edge Computing Environments: A Systematic Review and Performance Evaluation
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
von: Aribe Jr., Sales, et al.
Veröffentlicht: (2026)
Towards Verifiable Federated Unlearning: Framework, Challenges, and The Road Ahead
von: Nguyen, Thanh Linh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thanh Linh, et al.
Veröffentlicht: (2025)
Accelerating LLM Inference with Precomputed Query Storage
von: Park, Jay H., et al.
Veröffentlicht: (2025)
von: Park, Jay H., et al.
Veröffentlicht: (2025)
Beyond the Buzz: A Pragmatic Take on Inference Disaggregation
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
von: Mitra, Tiyasa, et al.
Veröffentlicht: (2025)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
von: Hwang, Changho, et al.
Veröffentlicht: (2025)
LLM Inference Serving: Survey of Recent Advances and Opportunities
von: Li, Baolin, et al.
Veröffentlicht: (2024)
von: Li, Baolin, et al.
Veröffentlicht: (2024)
Decentralized AI: Permissionless LLM Inference on POKT Network
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
von: Olshansky, Daniel, et al.
Veröffentlicht: (2024)
Large Language Model Partitioning for Low-Latency Inference at the Edge
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
von: Kafetzis, Dimitrios, et al.
Veröffentlicht: (2025)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
von: Lyu, Hongtao, et al.
Veröffentlicht: (2025)
Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
von: Li, Rongzhi, et al.
Veröffentlicht: (2025)
Opara: Exploiting Operator Parallelism for Expediting DNN Inference on GPUs
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
von: Chen, Aodong, et al.
Veröffentlicht: (2023)
A-IO: Adaptive Inference Orchestration for Memory-Bound NPUs
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
von: Zhang, Chen, et al.
Veröffentlicht: (2026)
Counting Without Running: Evaluating LLMs' Reasoning About Code Complexity
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
von: Bolet, Gregory, et al.
Veröffentlicht: (2025)
Seesaw: High-throughput LLM Inference via Model Re-sharding
von: Su, Qidong, et al.
Veröffentlicht: (2025)
von: Su, Qidong, et al.
Veröffentlicht: (2025)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
von: Wu, Linyu, et al.
Veröffentlicht: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
von: Stojkovic, Jovan, et al.
Veröffentlicht: (2025)
ELANA: A Simple Energy and Latency Analyzer for LLMs
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
von: Chiang, Hung-Yueh, et al.
Veröffentlicht: (2025)
Balanced and Elastic End-to-end Training of Dynamic LLMs
von: Wahib, Mohamed, et al.
Veröffentlicht: (2025)
von: Wahib, Mohamed, et al.
Veröffentlicht: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
von: Chen, Huamin, et al.
Veröffentlicht: (2026)
Dooly: Configuration-Agnostic, Redundancy-Aware Profiling for LLM Inference Simulation
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
von: Kim, Joon Ha, et al.
Veröffentlicht: (2026)
Understand and Accelerate Memory Processing Pipeline for Large Language Model Inference
von: He, Zifan, et al.
Veröffentlicht: (2026)
von: He, Zifan, et al.
Veröffentlicht: (2026)
Profiling-Driven Adaptive Distributed Transformer Inference on Embedded Edge Deployment
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
von: Qazi, Muhammad Azlan, et al.
Veröffentlicht: (2026)
Identifying and Mitigating Systemic Measurement Bias in Production LLM Inference Benchmarks
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
von: Chandrasekar, Ashok, et al.
Veröffentlicht: (2026)
Why Smaller Is Slower? Dimensional Misalignment in Compressed LLMs
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
von: Xin, Jihao, et al.
Veröffentlicht: (2026)
Cloud to Edge: Benchmarking LLM Inference On Hardware-Accelerated Single-Board Computers
von: Renney, Harri, et al.
Veröffentlicht: (2026)
von: Renney, Harri, et al.
Veröffentlicht: (2026)
KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
von: Zhang, Huawei, et al.
Veröffentlicht: (2025)
Reconstruction-Based Adaptive Scheduling Using AI Inferences in Safety-Critical Systems
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
von: Alshaer, Samer, et al.
Veröffentlicht: (2025)
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
von: Xu, Jiaming, et al.
Veröffentlicht: (2025)
FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference
von: Liu, Xing, et al.
Veröffentlicht: (2025)
von: Liu, Xing, et al.
Veröffentlicht: (2025)
AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
von: The AIBrix Team, et al.
Veröffentlicht: (2025)
Efficient MoE Inference with Fine-Grained Scheduling of Disaggregated Expert Parallelism
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
von: Pan, Xinglin, et al.
Veröffentlicht: (2025)
Verify Distributed Deep Learning Model Implementation Refinement with Iterative Relation Inference
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
von: Wang, Zhanghan, et al.
Veröffentlicht: (2025)
ECCENTRIC: Edge-Cloud Collaboration Framework for Distributed Inference Using Knowledge Adaptation
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Kamani, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
SparOA: Sparse and Operator-aware Hybrid Scheduling for Edge DNN Inference
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
von: Song, Jingwei, et al.
Veröffentlicht: (2025) -
Remoe: Towards Efficient and Low-Cost MoE Inference in Serverless Computing
von: Liu, Wentao, et al.
Veröffentlicht: (2025) -
Striking the Right Balance between Compute and Copy: Improving LLM Inferencing Under Speculative Decoding
von: Ramachandran, Arun, et al.
Veröffentlicht: (2025) -
Trust-Aware Routing for Distributed Generative AI Inference at the Edge
von: Nguyen, Chanh, et al.
Veröffentlicht: (2026) -
LLM as HPC Expert: Extending RAG Architecture for HPC Data
von: Miyashita, Yusuke, et al.
Veröffentlicht: (2024)