Serve Programs, Not Prompts
Fuente:
arXiv
Saved in:
| Main Authors: | Gim, In, Zhong, Lin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pie: A Programmable Serving System for Emerging LLM Applications
by: Gim, In, et al.
Published: (2025)
by: Gim, In, et al.
Published: (2025)
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024)
by: Li, Caihua, et al.
Published: (2024)
Cacheback: Speculative Decoding With Nothing But Cache
by: Ma, Zhiyao, et al.
Published: (2025)
by: Ma, Zhiyao, et al.
Published: (2025)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023)
by: Gim, In, et al.
Published: (2023)
Asynchronous LLM Function Calling
by: Gim, In, et al.
Published: (2024)
by: Gim, In, et al.
Published: (2024)
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
by: Wang, Ying, et al.
Published: (2025)
by: Wang, Ying, et al.
Published: (2025)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
by: Dong, Honghua, et al.
Published: (2024)
by: Dong, Honghua, et al.
Published: (2024)
Composing Policy Gradients and Prompt Optimization for Language Model Programs
by: Ziems, Noah, et al.
Published: (2025)
by: Ziems, Noah, et al.
Published: (2025)
Exploring Hybrid Question Answering via Program-based Prompting
by: Shi, Qi, et al.
Published: (2024)
by: Shi, Qi, et al.
Published: (2024)
From Data to Insights: Exploring Program-of-Thoughts Prompting for Chart Summarization
by: Qu, Yutong, et al.
Published: (2026)
by: Qu, Yutong, et al.
Published: (2026)
EFIM: Efficient Serving of LLMs for Infilling Tasks with Improved KV Cache Reuse
by: Guo, Tianyu, et al.
Published: (2025)
by: Guo, Tianyu, et al.
Published: (2025)
Cultural Value Differences of LLMs: Prompt, Language, and Model Size
by: Zhong, Qishuai, et al.
Published: (2024)
by: Zhong, Qishuai, et al.
Published: (2024)
Can LLM Prompting Serve as a Proxy for Static Analysis in Vulnerability Detection
by: Ceka, Ira, et al.
Published: (2024)
by: Ceka, Ira, et al.
Published: (2024)
PromptPrism: A Linguistically-Inspired Taxonomy for Prompts
by: Jeoung, Sullam, et al.
Published: (2025)
by: Jeoung, Sullam, et al.
Published: (2025)
P/D-Serve: Serving Disaggregated Large Language Model at Scale
by: Jin, Yibo, et al.
Published: (2024)
by: Jin, Yibo, et al.
Published: (2024)
Adaptive Prompt Structure Factorization: A Framework for Self-Discovering and Optimizing Compositional Prompt Programs
by: Liu, Haoyue, et al.
Published: (2026)
by: Liu, Haoyue, et al.
Published: (2026)
Prompt Programming for Cultural Bias and Alignment of Large Language Models
by: Eren, Maksim, et al.
Published: (2026)
by: Eren, Maksim, et al.
Published: (2026)
PromptExp: Multi-granularity Prompt Explanation of Large Language Models
by: Dong, Ximing, et al.
Published: (2024)
by: Dong, Ximing, et al.
Published: (2024)
Learning Task Decomposition to Assist Humans in Competitive Programming
by: Wen, Jiaxin, et al.
Published: (2024)
by: Wen, Jiaxin, et al.
Published: (2024)
BatchPrompt: Accomplish more with less
by: Lin, Jianzhe, et al.
Published: (2023)
by: Lin, Jianzhe, et al.
Published: (2023)
PANDA: Prompt Transfer Meets Knowledge Distillation for Efficient Model Adaptation
by: Zhong, Qihuang, et al.
Published: (2022)
by: Zhong, Qihuang, et al.
Published: (2022)
GAP: Graph-Assisted Prompts for Dialogue-based Medication Recommendation
by: Zhong, Jialun, et al.
Published: (2025)
by: Zhong, Jialun, et al.
Published: (2025)
Exchange of Perspective Prompting Enhances Reasoning in Large Language Models
by: Sun, Lin, et al.
Published: (2025)
by: Sun, Lin, et al.
Published: (2025)
AdaServe: Accelerating Multi-SLO LLM Serving with SLO-Customized Speculative Decoding
by: Li, Zikun, et al.
Published: (2025)
by: Li, Zikun, et al.
Published: (2025)
PromptMRG: Diagnosis-Driven Prompts for Medical Report Generation
by: Jin, Haibo, et al.
Published: (2023)
by: Jin, Haibo, et al.
Published: (2023)
Towards Resiliency in Large Language Model Serving with KevlarFlow
by: Qian, Shangshu, et al.
Published: (2026)
by: Qian, Shangshu, et al.
Published: (2026)
Safe to Serve: Aligning Instruction-Tuned Models for Safety and Helpfulness
by: Amballa, Avinash, et al.
Published: (2024)
by: Amballa, Avinash, et al.
Published: (2024)
Towards Pareto Optimal Throughput in Small Language Model Serving
by: Recasens, Pol G., et al.
Published: (2024)
by: Recasens, Pol G., et al.
Published: (2024)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
by: Chen, Xuanzhong, et al.
Published: (2024)
by: Chen, Xuanzhong, et al.
Published: (2024)
Optimizing Class-Level Probability Reweighting Coefficients for Equitable Prompting Accuracy
by: Lin, Ruixi, et al.
Published: (2024)
by: Lin, Ruixi, et al.
Published: (2024)
Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learning
by: Liu, Wenjin, et al.
Published: (2025)
by: Liu, Wenjin, et al.
Published: (2025)
The Impact of Prompt Programming on Function-Level Code Generation
by: Khojah, Ranim, et al.
Published: (2024)
by: Khojah, Ranim, et al.
Published: (2024)
Ensemble Debiasing Across Class and Sample Levels for Fairer Prompting Accuracy
by: Lin, Ruixi, et al.
Published: (2025)
by: Lin, Ruixi, et al.
Published: (2025)
P3: Prompts Promote Prompting
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
by: Luo, Run, et al.
Published: (2024)
by: Luo, Run, et al.
Published: (2024)
LangGPT: Rethinking Structured Reusable Prompt Design Framework for LLMs from the Programming Language
by: Wang, Ming, et al.
Published: (2024)
by: Wang, Ming, et al.
Published: (2024)
ROSE Doesn't Do That: Boosting the Safety of Instruction-Tuned Large Language Models with Reverse Prompt Contrastive Decoding
by: Zhong, Qihuang, et al.
Published: (2024)
by: Zhong, Qihuang, et al.
Published: (2024)
NavRAG: Generating User Demand Instructions for Embodied Navigation through Retrieval-Augmented LLM
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Similar Items
-
Pie: A Programmable Serving System for Emerging LLM Applications
by: Gim, In, et al.
Published: (2025) -
Confidential Prompting: Privacy-preserving LLM Inference on Cloud
by: Li, Caihua, et al.
Published: (2024) -
Cacheback: Speculative Decoding With Nothing But Cache
by: Ma, Zhiyao, et al.
Published: (2025) -
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
by: Gim, In, et al.
Published: (2023) -
Asynchronous LLM Function Calling
by: Gim, In, et al.
Published: (2024)