Pay for Hints, Not Answers: LLM Shepherding for Cost-Efficient Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Ziming, Sharma, Hardik, O'Toole, Evan, Champati, Jaya Prakash, Wu, Kui |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Regret and Low-Complexity Learning for Hierarchical Inference
by: Chattopadhyay, Sameep, et al.
Published: (2025)
by: Chattopadhyay, Sameep, et al.
Published: (2025)
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025)
by: Behera, Adarsh Prasad, et al.
Published: (2025)
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
Honest and Reliable Evaluation and Expert Equivalence Testing of Automated Neonatal Seizure Detection
by: Kljajic, Jovana, et al.
Published: (2025)
by: Kljajic, Jovana, et al.
Published: (2025)
Optimal Bayesian Stopping for Efficient Inference of Consistent LLM Answers
by: Huang, Jingkai, et al.
Published: (2026)
by: Huang, Jingkai, et al.
Published: (2026)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
by: Wang, Shengjie, et al.
Published: (2026)
by: Wang, Shengjie, et al.
Published: (2026)
Lexical Hints of Accuracy in LLM Reasoning Chains
by: Vanhoyweghen, Arne, et al.
Published: (2025)
by: Vanhoyweghen, Arne, et al.
Published: (2025)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
by: An, Zihao, et al.
Published: (2025)
by: An, Zihao, et al.
Published: (2025)
Minimizing Age of Detection for a Markov Source over a Lossy Channel
by: Garde, Shivang, et al.
Published: (2025)
by: Garde, Shivang, et al.
Published: (2025)
Improved Decision Module Selection for Hierarchical Inference in Resource-Constrained Edge Devices
by: Behera, Adarsh Prasad, et al.
Published: (2024)
by: Behera, Adarsh Prasad, et al.
Published: (2024)
Machine-learning competition to grade EEG background patterns in newborns with hypoxic-ischaemic encephalopathy
by: Magarelli, Fabio, et al.
Published: (2025)
by: Magarelli, Fabio, et al.
Published: (2025)
Point of View: Academic Librarians as STEM Retention Partners
by: O'Toole, Erin M.
Published: (2017)
by: O'Toole, Erin M.
Published: (2017)
Students, Artists, Television and Art Education.
by: O'Toole, Ora M.
Published: (1981)
by: O'Toole, Ora M.
Published: (1981)
Don Carlos Chimo del Perú: ¿del Común o cacique?
by: Rachel Sarah O'Toole
Published: (2011)
by: Rachel Sarah O'Toole
Published: (2011)
Pie: Pooling CPU Memory for LLM Inference
by: Xu, Yi, et al.
Published: (2024)
by: Xu, Yi, et al.
Published: (2024)
Anytime-Valid Answer Sufficiency Certificates for LLM Generation via Sequential Information Lift
by: Akter, Sanjeda, et al.
Published: (2025)
by: Akter, Sanjeda, et al.
Published: (2025)
Prompt Smart, Pay Less: Cost-Aware APO for Real-World Applications
by: Choudhari, Jayesh, et al.
Published: (2025)
by: Choudhari, Jayesh, et al.
Published: (2025)
Unconstrained Body Recognition at Altitude and Range: Comparing Four Approaches
by: Myers, Blake A, et al.
Published: (2025)
by: Myers, Blake A, et al.
Published: (2025)
HCRMP: A LLM-Hinted Contextual Reinforcement Learning Framework for Autonomous Driving
by: Chen, Zhiwen, et al.
Published: (2025)
by: Chen, Zhiwen, et al.
Published: (2025)
Scaling convolutional neural networks achieves expert-level seizure detection in neonatal EEG
by: Hogan, Robert, et al.
Published: (2024)
by: Hogan, Robert, et al.
Published: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
by: Ning, Lin, et al.
Published: (2024)
by: Ning, Lin, et al.
Published: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
by: Zheng, Haizhong, et al.
Published: (2025)
by: Zheng, Haizhong, et al.
Published: (2025)
Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization
by: Zhang, Qingyang, et al.
Published: (2025)
by: Zhang, Qingyang, et al.
Published: (2025)
Federated Learning of Binary Neural Networks: Enabling Low-Cost Inference
by: Shankar, Nitin Priyadarshini, et al.
Published: (2026)
by: Shankar, Nitin Priyadarshini, et al.
Published: (2026)
Cost-Efficient Multimodal LLM Inference via Cross-Tier GPU Heterogeneity
by: Yu, Donglin
Published: (2026)
by: Yu, Donglin
Published: (2026)
CHAI: Clustered Head Attention for Efficient LLM Inference
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing
by: Li, Yang
Published: (2025)
by: Li, Yang
Published: (2025)
SparQ Attention: Bandwidth-Efficient LLM Inference
by: Ribar, Luka, et al.
Published: (2023)
by: Ribar, Luka, et al.
Published: (2023)
Self-Hinting Language Models Enhance Reinforcement Learning
by: Liao, Baohao, et al.
Published: (2026)
by: Liao, Baohao, et al.
Published: (2026)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
COLA: Continual Learning via Autoencoder Retrieval of Adapters
by: Mandivarapu, Jaya Krishna
Published: (2025)
by: Mandivarapu, Jaya Krishna
Published: (2025)
StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason
by: Zhang, Kaiyi, et al.
Published: (2025)
by: Zhang, Kaiyi, et al.
Published: (2025)
Cost and Reward Infused Metric Elicitation
by: Bhateja, Chethan, et al.
Published: (2025)
by: Bhateja, Chethan, et al.
Published: (2025)
Flow-Controlled Scheduling for LLM Inference with Provable Stability Guarantees
by: Dong, Zhuolun, et al.
Published: (2026)
by: Dong, Zhuolun, et al.
Published: (2026)
LLMEasyQuant: Scalable Quantization for Parallel and Distributed LLM Inference
by: Liu, Dong, et al.
Published: (2024)
by: Liu, Dong, et al.
Published: (2024)
Efficient Mixture-of-Experts LLM Inference with Apple Silicon NPUs
by: Benazir, Afsara, et al.
Published: (2026)
by: Benazir, Afsara, et al.
Published: (2026)
Eagle: Efficient Training-Free Router for Multi-LLM Inference
by: Zhao, Zesen, et al.
Published: (2024)
by: Zhao, Zesen, et al.
Published: (2024)
LLM Cannot Discover Causality, and Should Be Restricted to Non-Decisional Support in Causal Discovery
by: Wu, Xingyu, et al.
Published: (2025)
by: Wu, Xingyu, et al.
Published: (2025)
Similar Items
-
Low-Regret and Low-Complexity Learning for Hierarchical Inference
by: Chattopadhyay, Sameep, et al.
Published: (2025) -
Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques
by: Behera, Adarsh Prasad, et al.
Published: (2025) -
Online Algorithms for Hierarchical Inference in Deep Learning applications at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2023) -
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025) -
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
by: Behera, Adarsh Prasad, et al.
Published: (2024)