Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zonghang, Li, Tao, Feng, Wenjiao, Xiao, Rongxing, She, Jianshu, Huang, Hong, Guizani, Mohsen, Yu, Hongfang, Ho, Qirong, Xiang, Wei, Liu, Steve |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
by: Feng, Wenjiao, et al.
Published: (2025)
by: Feng, Wenjiao, et al.
Published: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024)
by: Li, Zonghang, et al.
Published: (2024)
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026)
by: Wang, Yuchen, et al.
Published: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
Redefining Clustered Federated Learning for System Identification: The Path of ClusterCraft
by: Keçeci, Ertuğrul, et al.
Published: (2025)
by: Keçeci, Ertuğrul, et al.
Published: (2025)
Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections
by: Song, Wenzhe, et al.
Published: (2026)
by: Song, Wenzhe, et al.
Published: (2026)
Go Big or Go Home: Simulating Mobbing Behavior with Braitenbergian Robots
by: Sanoubari, Elaheh
Published: (2026)
by: Sanoubari, Elaheh
Published: (2026)
FSFM: A Biologically-Inspired Framework for Selective Forgetting of Agent Memory
by: Gu, Yingjie, et al.
Published: (2026)
by: Gu, Yingjie, et al.
Published: (2026)
QTypeMix: Enhancing Multi-Agent Cooperative Strategies through Heterogeneous and Homogeneous Value Decomposition
by: Fu, Songchen, et al.
Published: (2024)
by: Fu, Songchen, et al.
Published: (2024)
Cooperative and Asynchronous Transformer-based Mission Planning for Heterogeneous Teams of Mobile Robots
by: Farjadnasab, Milad, et al.
Published: (2024)
by: Farjadnasab, Milad, et al.
Published: (2024)
AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation
by: Costa, Igor
Published: (2026)
by: Costa, Igor
Published: (2026)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
by: Chen, Wen-Tse, et al.
Published: (2024)
by: Chen, Wen-Tse, et al.
Published: (2024)
Semantic Risk-Aware Heuristic Planning for Robotic Navigation in Dynamic Environments: An LLM-Inspired Approach
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
by: Durrani, Hamza Ahmed, et al.
Published: (2026)
FORMICA: Decision-Focused Learning for Communication-Free Multi-Robot Task Allocation
by: Lopez, Antonio, et al.
Published: (2026)
by: Lopez, Antonio, et al.
Published: (2026)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
by: Hu, Pan
Published: (2025)
by: Hu, Pan
Published: (2025)
Beyond Rating: A Comprehensive Evaluation and Benchmark for AI Reviews
by: Li, Bowen, et al.
Published: (2026)
by: Li, Bowen, et al.
Published: (2026)
Agent Capsules: Quality-Gated Granularity Control for Multi-Agent LLM Pipelines
by: Ray, Aninda
Published: (2026)
by: Ray, Aninda
Published: (2026)
PRIMA: Operational Patterns for Resilient Multi-Agent Research with Verifiable Identity and Convergent Feedback
by: Annapureddy, Sasank
Published: (2026)
by: Annapureddy, Sasank
Published: (2026)
Analysis of Total Variation Minimization for Clustered Federated Learning
by: Jung, A.
Published: (2024)
by: Jung, A.
Published: (2024)
MMiC: Mitigating Modality Incompleteness in Clustered Federated Learning
by: Yang, Lishan, et al.
Published: (2025)
by: Yang, Lishan, et al.
Published: (2025)
HULK: Large-scale Hierarchical Coordination under Continual and Uncertain Temporal Tasks
by: Luo, Qingyuan, et al.
Published: (2026)
by: Luo, Qingyuan, et al.
Published: (2026)
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
by: Liu, Dongxiao, et al.
Published: (2026)
by: Liu, Dongxiao, et al.
Published: (2026)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
by: Alanazi, Ahmed, et al.
Published: (2025)
by: Alanazi, Ahmed, et al.
Published: (2025)
Enhancing Heterogeneous Multi-Agent Cooperation in Decentralized MARL via GNN-driven Intrinsic Rewards
by: Monon, Jahir Sadik, et al.
Published: (2024)
by: Monon, Jahir Sadik, et al.
Published: (2024)
Exploring Robust Multi-Agent Workflows for Environmental Data Management
by: Guan, Boyuan, et al.
Published: (2026)
by: Guan, Boyuan, et al.
Published: (2026)
Federated Single-Agent Robotics: Multi-Robot Coordination Without Intra-Robot Multi-Agent Fragmentation
by: Qin, Xue, et al.
Published: (2026)
by: Qin, Xue, et al.
Published: (2026)
EmbodiedGovBench: A Benchmark for Governance, Recovery, and Upgrade Safety in Embodied Agent Systems
by: Qin, Xue, et al.
Published: (2026)
by: Qin, Xue, et al.
Published: (2026)
Instruction-Level Weight Shaping: A Framework for Self-Improving AI Agents
by: Costa, Rimom
Published: (2025)
by: Costa, Rimom
Published: (2025)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
by: Rossi, Maximillian, et al.
Published: (2026)
by: Rossi, Maximillian, et al.
Published: (2026)
A Super-Learner with Large Language Models for Medical Emergency Advising
by: Aityan, Sergey K., et al.
Published: (2025)
by: Aityan, Sergey K., et al.
Published: (2025)
PRISM-Consult: A Panel-of-Experts Architecture for Clinician-Aligned Diagnosis
by: Levine, Lionel, et al.
Published: (2025)
by: Levine, Lionel, et al.
Published: (2025)
One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents
by: Hong, Yoosung
Published: (2026)
by: Hong, Yoosung
Published: (2026)
Advancing Multimodal Agent Reasoning with Long-Term Neuro-Symbolic Memory
by: Jiang, Rongjie, et al.
Published: (2026)
by: Jiang, Rongjie, et al.
Published: (2026)
Applying Cognitive Design Patterns to General LLM Agents
by: Wray, Robert E., et al.
Published: (2025)
by: Wray, Robert E., et al.
Published: (2025)
Privacy Preserving Multi Agent Path Finding
by: Lehman, Rotem Lev, et al.
Published: (2026)
by: Lehman, Rotem Lev, et al.
Published: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Collaborative On-Sensor Array Cameras
by: Sun, Jipeng, et al.
Published: (2025)
by: Sun, Jipeng, et al.
Published: (2025)
Similar Items
-
TPI-LLM: Serving 70B-scale LLMs Efficiently on Low-resource Edge Devices
by: Li, Zonghang, et al.
Published: (2024) -
Learning In Chaos: Efficient Autoscaling and Self-Healing for Multi-Party Distributed Training
by: Feng, Wenjiao, et al.
Published: (2025) -
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026) -
Accelerating Geo-distributed Machine Learning with Network-Aware Adaptive Tree and Auxiliary Route
by: Li, Zonghang, et al.
Published: (2024) -
PARNESS: A Paper Harness for End-to-End Automated Scientific Research with Dynamic Workflows, Full-Text Indexing, and Cross-Run Knowledge Accumulation
by: Wang, Yuchen, et al.
Published: (2026)