Viability and Performance of a Private LLM Server for SMBs: A Benchmark Analysis of Qwen3-30B on Consumer-Grade Hardware
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khalil, Alex, Heilles, Guillaume, Parraga, Maria, Heilles, Simon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Experimental Analysis of Server-Side Caching for Web Performance
von: Umar, Mohammad, et al.
Veröffentlicht: (2026)
von: Umar, Mohammad, et al.
Veröffentlicht: (2026)
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026)
Experiences with Model Context Protocol Servers for Science and High Performance Computing
von: Pan, Haochen, et al.
Veröffentlicht: (2025)
von: Pan, Haochen, et al.
Veröffentlicht: (2025)
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
von: Ji, Shixin, et al.
Veröffentlicht: (2024)
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
von: Li, Zhilin, et al.
Veröffentlicht: (2025)
Analysis of Server Throughput For Managed Big Data Analytics Frameworks
von: Anagnostakis, Emmanouil, et al.
Veröffentlicht: (2025)
von: Anagnostakis, Emmanouil, et al.
Veröffentlicht: (2025)
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
Benchmarking Compound AI Applications for Hardware-Software Co-Design
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
von: Samuthrsindh, Paramuth, et al.
Veröffentlicht: (2026)
Are Bus-Mounted Edge Servers Feasible?
von: Li, Xuezhi, et al.
Veröffentlicht: (2025)
von: Li, Xuezhi, et al.
Veröffentlicht: (2025)
AgentServe: Algorithm-System Co-Design for Efficient Agentic AI Serving on a Consumer-Grade GPU
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
von: Zhang, Yuning, et al.
Veröffentlicht: (2026)
Gateways for Institutional-Grade Commerce and Interoperability of Digital Assets
von: Belchior, Rafael, et al.
Veröffentlicht: (2025)
von: Belchior, Rafael, et al.
Veröffentlicht: (2025)
Practical Federated Learning without a Server
von: Dhasade, Akash, et al.
Veröffentlicht: (2025)
von: Dhasade, Akash, et al.
Veröffentlicht: (2025)
Exploring the Viability of Unikernels for ARM-powered Edge Computing
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
von: Kaiser, Shahidullah, et al.
Veröffentlicht: (2024)
Revisiting Parameter Server in LLM Post-Training
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
von: Wan, Xinyi, et al.
Veröffentlicht: (2026)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
von: Shubham, et al.
Veröffentlicht: (2024)
von: Shubham, et al.
Veröffentlicht: (2024)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
von: Ovi, Md Sultanul Islam
Veröffentlicht: (2025)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
A Multi-Server Information-Sharing Environment for Cross-Party Collaboration on A Private Cloud
von: Zhang, Jianping, et al.
Veröffentlicht: (2024)
von: Zhang, Jianping, et al.
Veröffentlicht: (2024)
Predictive Analysis of CFPB Consumer Complaints Using Machine Learning
von: Vaishnav, Dhwani, et al.
Veröffentlicht: (2024)
von: Vaishnav, Dhwani, et al.
Veröffentlicht: (2024)
Beyond End-to-End: Dynamic Chain Optimization for Private LLM Adaptation on the Edge
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
von: Wu, Yebo, et al.
Veröffentlicht: (2026)
6G EdgeAI: Performance Evaluation and Analysis
von: Yang, Chien-Sheng, et al.
Veröffentlicht: (2025)
von: Yang, Chien-Sheng, et al.
Veröffentlicht: (2025)
Adversarial Analysis of the Differentially-Private Federated Learning in Cyber-Physical Critical Infrastructures
von: Hossain, Md Tamjid, et al.
Veröffentlicht: (2022)
von: Hossain, Md Tamjid, et al.
Veröffentlicht: (2022)
ReviveMoE: Fast Recovery for Hardware Failures in Large-Scale MoE LLM Inference Deployments
von: Li, Haley, et al.
Veröffentlicht: (2026)
von: Li, Haley, et al.
Veröffentlicht: (2026)
DiSCo: Device-Server Collaborative LLM-Based Text Streaming Services
von: Sun, Ting, et al.
Veröffentlicht: (2025)
von: Sun, Ting, et al.
Veröffentlicht: (2025)
Gaia: Hybrid Hardware Acceleration for Serverless AI in the 3D Compute Continuum
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
von: Reisecker, Maximilian, et al.
Veröffentlicht: (2025)
More is Different: Prototyping and Analyzing a New Form of Edge Server with Massive Mobile SoCs
von: Zhang, Li, et al.
Veröffentlicht: (2022)
von: Zhang, Li, et al.
Veröffentlicht: (2022)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
von: Duan, Tao, et al.
Veröffentlicht: (2025)
von: Duan, Tao, et al.
Veröffentlicht: (2025)
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
von: Li, Shengwei, et al.
Veröffentlicht: (2023)
Understanding the Performance and Power of LLM Inferencing on Edge Accelerators
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
von: Arya, Mayank, et al.
Veröffentlicht: (2025)
LLM & HPC:Benchmarking DeepSeek's Performance in High-Performance Computing Tasks
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
von: Nader, Noujoud, et al.
Veröffentlicht: (2025)
Benchmarking the Performance of Large Language Models on the Cerebras Wafer Scale Engine
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
von: Zhang, Zuoning, et al.
Veröffentlicht: (2024)
Towards Energy-Efficient Serverless Computing with Hardware Isolation
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
von: Carl, Natalie, et al.
Veröffentlicht: (2025)
SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
von: Cheng, Ke, et al.
Veröffentlicht: (2024)
Benchmarking Message Brokers for IoT Edge Computing: A Comprehensive Performance Study
von: Paul, Tapajit Chandra, et al.
Veröffentlicht: (2026)
von: Paul, Tapajit Chandra, et al.
Veröffentlicht: (2026)
Performance Evaluation of Hashing Algorithms on Commodity Hardware
von: Pandya, Marut
Veröffentlicht: (2024)
von: Pandya, Marut
Veröffentlicht: (2024)
IM-PIR: In-Memory Private Information Retrieval
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2025)
von: Mwaisela, Mpoki, et al.
Veröffentlicht: (2025)
Availability Modeling for Blockchain Provisioning in Private Clouds
von: Dantas, J, et al.
Veröffentlicht: (2025)
von: Dantas, J, et al.
Veröffentlicht: (2025)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
von: Thüring, Tim, et al.
Veröffentlicht: (2026)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
von: Islam, Tanzima Z., et al.
Veröffentlicht: (2024)
Extracting the Potential of Emerging Hardware Accelerators for Symmetric Eigenvalue Decomposition
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
von: Wang, Hansheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Experimental Analysis of Server-Side Caching for Web Performance
von: Umar, Mohammad, et al.
Veröffentlicht: (2026) -
PipeMax: Enhancing Offline LLM Inference on Commodity GPU Servers
von: Zhang, Hongbin, et al.
Veröffentlicht: (2026) -
Experiences with Model Context Protocol Servers for Science and High Performance Computing
von: Pan, Haochen, et al.
Veröffentlicht: (2025) -
SCARIF: Towards Carbon Modeling of Cloud Servers with Accelerators
von: Ji, Shixin, et al.
Veröffentlicht: (2024) -
TailBench++: Flexible Multi-Client, Multi-Server Benchmarking for Latency-Critical Workloads
von: Li, Zhilin, et al.
Veröffentlicht: (2025)