Inference-Time Distillation: Cost-Efficient Agents Without Fine-Tuning or Manual Prompt Engineering
Fuente:
arXiv
Saved in:
| Main Authors: | Sarukkai, Vishnu, Gupta, Asanshay, Hong, James, Gharbi, Michaël, Fatahalian, Kayvon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
by: Sarukkai, Vishnu, et al.
Published: (2025)
by: Sarukkai, Vishnu, et al.
Published: (2025)
Automated Rewards via LLM-Generated Progress Functions
by: Sarukkai, Vishnu, et al.
Published: (2024)
by: Sarukkai, Vishnu, et al.
Published: (2024)
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
by: Xu, Pei, et al.
Published: (2025)
by: Xu, Pei, et al.
Published: (2025)
Block and Detail: Scaffolding Sketch-to-Image Generation
by: Sarukkai, Vishnu, et al.
Published: (2024)
by: Sarukkai, Vishnu, et al.
Published: (2024)
Learning to Move Like Professional Counter-Strike Players
by: Durst, David, et al.
Published: (2024)
by: Durst, David, et al.
Published: (2024)
Learning Subject-Aware Cropping by Outpainting Professional Photos
by: Hong, James, et al.
Published: (2023)
by: Hong, James, et al.
Published: (2023)
AI Metropolis: Scaling Large Language Model-based Multi-Agent Simulation with Out-of-order Execution
by: Xie, Zhiqiang, et al.
Published: (2024)
by: Xie, Zhiqiang, et al.
Published: (2024)
Beyond LoRA: Exploring Efficient Fine-Tuning Techniques for Time Series Foundational Models
by: Gupta, Divij, et al.
Published: (2024)
by: Gupta, Divij, et al.
Published: (2024)
Symbiosis: Multi-Adapter Inference and Fine-Tuning
by: Gupta, Saransh, et al.
Published: (2025)
by: Gupta, Saransh, et al.
Published: (2025)
Memory-Space Visual Prompting for Efficient Vision-Language Fine-Tuning
by: Jie, Shibo, et al.
Published: (2024)
by: Jie, Shibo, et al.
Published: (2024)
Order-Independence Without Fine Tuning
by: McIlroy-Young, Reid, et al.
Published: (2024)
by: McIlroy-Young, Reid, et al.
Published: (2024)
AgentFlux: Decoupled Fine-Tuning & Inference for On-Device Agentic Systems
by: Kadekodi, Rohan, et al.
Published: (2025)
by: Kadekodi, Rohan, et al.
Published: (2025)
IPGO: Indirect Prompt Gradient Optimization for Parameter-Efficient Prompt-level Fine-Tuning on Text-to-Image Models
by: Ye, Jianping, et al.
Published: (2025)
by: Ye, Jianping, et al.
Published: (2025)
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025)
by: Bhathal, Tanvir, et al.
Published: (2025)
Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
Eliciting Fine-Tuned Transformer Capabilities via Inference-Time Techniques
by: Sharma, Asankhaya
Published: (2025)
by: Sharma, Asankhaya
Published: (2025)
Robust Graph Fine-Tuning with Adversarial Graph Prompting
by: Zhang, Ziyan, et al.
Published: (2026)
by: Zhang, Ziyan, et al.
Published: (2026)
PromptIntern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Language Model Fine-tuning
by: Zou, Jiaru, et al.
Published: (2024)
by: Zou, Jiaru, et al.
Published: (2024)
PromptRL: Prompt Matters in RL for Flow-Based Image Generation
by: Wang, Fu-Yun, et al.
Published: (2026)
by: Wang, Fu-Yun, et al.
Published: (2026)
FedHPL: Efficient Heterogeneous Federated Learning with Prompt Tuning and Logit Distillation
by: Ma, Yuting, et al.
Published: (2024)
by: Ma, Yuting, et al.
Published: (2024)
Modular Multimodal Classification Without Fine-Tuning: A Simple Compositional Approach
by: Bergström, Herman, et al.
Published: (2026)
by: Bergström, Herman, et al.
Published: (2026)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
MELINOE: Fine-Tuning Enables Memory-Efficient Inference for Mixture-of-Experts Models
by: Raje, Arian, et al.
Published: (2026)
by: Raje, Arian, et al.
Published: (2026)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
by: Li, Zelin, et al.
Published: (2026)
by: Li, Zelin, et al.
Published: (2026)
Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs
by: Hübotter, Jonas, et al.
Published: (2024)
by: Hübotter, Jonas, et al.
Published: (2024)
DynaPrompt: Dynamic Test-Time Prompt Tuning
by: Xiao, Zehao, et al.
Published: (2025)
by: Xiao, Zehao, et al.
Published: (2025)
TuneShift-KD: Knowledge Distillation and Transfer for Fine-tuned Models
by: Guan, Yushi, et al.
Published: (2026)
by: Guan, Yushi, et al.
Published: (2026)
Prompt Tuning for Natural Language to SQL with Embedding Fine-Tuning and RAG
by: Jang, Jisoo, et al.
Published: (2025)
by: Jang, Jisoo, et al.
Published: (2025)
Communication-Efficient Federated Fine-Tuning
by: Theologitis, Michael, et al.
Published: (2025)
by: Theologitis, Michael, et al.
Published: (2025)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
by: Park, Juneyoung, et al.
Published: (2026)
by: Park, Juneyoung, et al.
Published: (2026)
Representation Without Reward: A JEPA Audit for LLM Fine-Tuning
by: Sengupta, Biswa
Published: (2026)
by: Sengupta, Biswa
Published: (2026)
Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion Parameters
by: Wang, Weiyang, et al.
Published: (2023)
by: Wang, Weiyang, et al.
Published: (2023)
Inference Offloading for Cost-Sensitive Binary Classification at the Edge
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
by: Moothedath, Vishnu Narayanan, et al.
Published: (2025)
Constrained Edge AI Deployment: Fine-Tuning vs Distillation for LLM Compression
by: Sander, Jacob, et al.
Published: (2025)
by: Sander, Jacob, et al.
Published: (2025)
Scaling Pretrained Representations Enables Label-Free Out-of-Distribution Detection Without Fine-Tuning
by: Barkley, Brett, et al.
Published: (2026)
by: Barkley, Brett, et al.
Published: (2026)
LoRA Fine-Tuning Without GPUs: A CPU-Efficient Meta-Generation Framework for LLMs
by: Arabpour, Reza, et al.
Published: (2025)
by: Arabpour, Reza, et al.
Published: (2025)
Quaff: Quantized Parameter-Efficient Fine-Tuning under Outlier Spatial Stability Hypothesis
by: Huang, Hong, et al.
Published: (2025)
by: Huang, Hong, et al.
Published: (2025)
Task-Adaptive Parameter-Efficient Fine-Tuning for Weather Foundation Models
by: Cao, Shilei, et al.
Published: (2025)
by: Cao, Shilei, et al.
Published: (2025)
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
Efficient Inference Using Large Language Models with Limited Human Data: Fine-Tuning then Rectification
by: Wang, Lei, et al.
Published: (2025)
by: Wang, Lei, et al.
Published: (2025)
Similar Items
-
Self-Generated In-Context Examples Improve LLM Agents for Sequential Decision-Making Tasks
by: Sarukkai, Vishnu, et al.
Published: (2025) -
Automated Rewards via LLM-Generated Progress Functions
by: Sarukkai, Vishnu, et al.
Published: (2024) -
Learning to Ball: Composing Policies for Long-Horizon Basketball Moves
by: Xu, Pei, et al.
Published: (2025) -
Block and Detail: Scaffolding Sketch-to-Image Generation
by: Sarukkai, Vishnu, et al.
Published: (2024) -
Learning to Move Like Professional Counter-Strike Players
by: Durst, David, et al.
Published: (2024)