Design and Evaluation of Cost-Aware PoQ for Decentralized LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Tian, Arther, Ding, Alex, Chen, Frank, Wu, Alan, Chan, Aaron, Zhang, Bruce |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
by: Tian, Arther, et al.
Published: (2026)
by: Tian, Arther, et al.
Published: (2026)
A Multi-Dimensional Quality Scoring Framework for Decentralized LLM Inference with Proof of Quality
by: Tian, Arther, et al.
Published: (2026)
by: Tian, Arther, et al.
Published: (2026)
Optimistic TEE-Rollups: A Hybrid Architecture for Scalable and Verifiable Generative AI Inference on Blockchain
by: Chan, Aaron, et al.
Published: (2025)
by: Chan, Aaron, et al.
Published: (2025)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026)
by: Ding, Wenxuan, et al.
Published: (2026)
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
CR^2: Cost-Aware Risk-Controlled Routing for Wireless Device-Edge LLM Inference
by: Xue, Nan, et al.
Published: (2026)
by: Xue, Nan, et al.
Published: (2026)
DeServe: Towards Affordable Offline LLM Inference via Decentralization
by: Wu, Linyu, et al.
Published: (2025)
by: Wu, Linyu, et al.
Published: (2025)
OCCAM: Towards Cost-Efficient and Accuracy-Aware Classification Inference
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost
by: Tan, Qitao, et al.
Published: (2025)
by: Tan, Qitao, et al.
Published: (2025)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
Polyrating: A Cost-Effective and Bias-Aware Rating System for LLM Evaluation
by: Dekoninck, Jasper, et al.
Published: (2024)
by: Dekoninck, Jasper, et al.
Published: (2024)
Cost-Aware Model Orchestration for LLM-based Systems
by: Smirnova, Daria, et al.
Published: (2025)
by: Smirnova, Daria, et al.
Published: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
by: Olshansky, Daniel, et al.
Published: (2024)
by: Olshansky, Daniel, et al.
Published: (2024)
Cost-Awareness in Tree-Search LLM Planning: A Systematic Study
by: Zhang, Zihao, et al.
Published: (2025)
by: Zhang, Zihao, et al.
Published: (2025)
Trust by Design: Skill Profiles for Transparent, Cost-Aware LLM Routing
by: Okamoto, Mika, et al.
Published: (2026)
by: Okamoto, Mika, et al.
Published: (2026)
CATP-LLM: Empowering Large Language Models for Cost-Aware Tool Planning
by: Wu, Duo, et al.
Published: (2024)
by: Wu, Duo, et al.
Published: (2024)
FairBatching: Fairness-Aware Batch Formation for LLM Inference
by: Lyu, Hongtao, et al.
Published: (2025)
by: Lyu, Hongtao, et al.
Published: (2025)
Speculative Decoding in Decentralized LLM Inference: Turning Communication Latency into Computation Throughput
by: Song, Jingwei, et al.
Published: (2025)
by: Song, Jingwei, et al.
Published: (2025)
Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing
by: Ding, Dujian, et al.
Published: (2024)
by: Ding, Dujian, et al.
Published: (2024)
Response-Aware User Memory Selection for LLM Personalization
by: Fisher, Jillian, et al.
Published: (2026)
by: Fisher, Jillian, et al.
Published: (2026)
ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation
by: Yuan, Boqin, et al.
Published: (2026)
by: Yuan, Boqin, et al.
Published: (2026)
RM-PoT: Reformulating Mathematical Problems and Solving via Program of Thoughts
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
by: Zhang, Hongbin, et al.
Published: (2025)
by: Zhang, Hongbin, et al.
Published: (2025)
eFedLLM: Efficient LLM Inference Based on Federated Learning
by: Ding, Shengwen, et al.
Published: (2024)
by: Ding, Shengwen, et al.
Published: (2024)
Distribution-Aware Algorithm Design with LLM Agents
by: Koganti, Saharsh, et al.
Published: (2026)
by: Koganti, Saharsh, et al.
Published: (2026)
Threshold-Based Exclusive Batching for LLM Inference
by: Zhang, Weifang, et al.
Published: (2026)
by: Zhang, Weifang, et al.
Published: (2026)
PoAct: Policy and Action Dual-Control Agent for Generalized Applications
by: Yuan, Guozhi, et al.
Published: (2025)
by: Yuan, Guozhi, et al.
Published: (2025)
Cost-optimal Sequential Testing via Doubly Robust Q-learning
by: Zhou, Doudou, et al.
Published: (2026)
by: Zhou, Doudou, et al.
Published: (2026)
GAR: Carbon-Aware Routing for LLM Inference via Constrained Optimization
by: Sheshanarayana, Disha, et al.
Published: (2026)
by: Sheshanarayana, Disha, et al.
Published: (2026)
LLM-Hanabi: Evaluating Multi-Agent Gameplays with Theory-of-Mind and Rationale Inference in Imperfect Information Collaboration Game
by: Liang, Fangzhou, et al.
Published: (2025)
by: Liang, Fangzhou, et al.
Published: (2025)
Towards Heterogeneity-Aware and Energy-Efficient Topology Optimization for Decentralized Federated Learning in Edge Environment
by: Liu, Yuze, et al.
Published: (2025)
by: Liu, Yuze, et al.
Published: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
by: Agarwal, Anisha, et al.
Published: (2024)
by: Agarwal, Anisha, et al.
Published: (2024)
Context-Aware Inference via Performance Forecasting in Decentralized Learning Networks
by: Pfeffer, Joel, et al.
Published: (2025)
by: Pfeffer, Joel, et al.
Published: (2025)
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings
by: Zhang, Yiqun, et al.
Published: (2026)
by: Zhang, Yiqun, et al.
Published: (2026)
Best-of-Q: Improving VLM agents with Q-function Action Ranking at Inference
by: Biré, Emilien, et al.
Published: (2026)
by: Biré, Emilien, et al.
Published: (2026)
On-Chain Decentralized Learning and Cost-Effective Inference for DeFi Attack Mitigation
by: Alhaidari, Abdulrahman, et al.
Published: (2025)
by: Alhaidari, Abdulrahman, et al.
Published: (2025)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
by: Lu, Guanxi, et al.
Published: (2025)
by: Lu, Guanxi, et al.
Published: (2025)
Topology-Aware Knowledge Propagation in Decentralized Learning
by: Sakarvadia, Mansi, et al.
Published: (2025)
by: Sakarvadia, Mansi, et al.
Published: (2025)
ConPoSe: LLM-Guided Contact Point Selection for Scalable Cooperative Object Pushing
by: Steinkrüger, Noah, et al.
Published: (2025)
by: Steinkrüger, Noah, et al.
Published: (2025)
Bridging Social Psychology and LLM Reasoning: Conflict-Aware Meta-Review Generation via Cognitive Alignment
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
Similar Items
-
Adaptive and Robust Cost-Aware Proof of Quality for Decentralized LLM Inference Networks
by: Tian, Arther, et al.
Published: (2026) -
A Multi-Dimensional Quality Scoring Framework for Decentralized LLM Inference with Proof of Quality
by: Tian, Arther, et al.
Published: (2026) -
Optimistic TEE-Rollups: A Hybrid Architecture for Scalable and Verifiable Generative AI Inference on Blockchain
by: Chan, Aaron, et al.
Published: (2025) -
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
by: Ding, Wenxuan, et al.
Published: (2026) -
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
by: Chen, Huamin, et al.
Published: (2026)