ICPC: In-context Prompt Compression with Faster Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Ziyang, Liu, Yuyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
di: Kummer, Cornelius, et al.
Pubblicazione: (2026)
di: Kummer, Cornelius, et al.
Pubblicazione: (2026)
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
di: Hu, Jinwu, et al.
Pubblicazione: (2025)
di: Hu, Jinwu, et al.
Pubblicazione: (2025)
EFPC: Towards Efficient and Flexible Prompt Compression
di: Cao, Yun-Hao, et al.
Pubblicazione: (2025)
di: Cao, Yun-Hao, et al.
Pubblicazione: (2025)
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
di: Han, Chao, et al.
Pubblicazione: (2025)
di: Han, Chao, et al.
Pubblicazione: (2025)
Discrete Prompt Compression with Reinforcement Learning
di: Jung, Hoyoun, et al.
Pubblicazione: (2023)
di: Jung, Hoyoun, et al.
Pubblicazione: (2023)
Provable Benefits of Task-Specific Prompts for In-context Learning
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
di: Chang, Xiangyu, et al.
Pubblicazione: (2025)
Learning to Compress Prompt in Natural Language Formats
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
di: Chuang, Yu-Neng, et al.
Pubblicazione: (2024)
Parse Trees Guided LLM Prompt Compression
di: Mao, Wenhao, et al.
Pubblicazione: (2024)
di: Mao, Wenhao, et al.
Pubblicazione: (2024)
Efficient and Effective Prompt Tuning via Prompt Decomposition and Compressed Outer Product
di: Lan, Pengxiang, et al.
Pubblicazione: (2025)
di: Lan, Pengxiang, et al.
Pubblicazione: (2025)
Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing
di: Liu, Ziyang
Pubblicazione: (2026)
di: Liu, Ziyang
Pubblicazione: (2026)
Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations
di: Liu, Ziyang
Pubblicazione: (2026)
di: Liu, Ziyang
Pubblicazione: (2026)
SCOPE: A Generative Approach for LLM Prompt Compression
di: Zhang, Tinghui, et al.
Pubblicazione: (2025)
di: Zhang, Tinghui, et al.
Pubblicazione: (2025)
An Empirical Study on Prompt Compression for Large Language Models
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
Prompt-SAW: Leveraging Relation-Aware Graphs for Textual Prompt Compression
di: Ali, Muhammad Asif, et al.
Pubblicazione: (2024)
di: Ali, Muhammad Asif, et al.
Pubblicazione: (2024)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
di: Adamska, Marta, et al.
Pubblicazione: (2025)
di: Adamska, Marta, et al.
Pubblicazione: (2025)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
di: Wang, Yuhao, et al.
Pubblicazione: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
Dynamic Depth Decoding: Faster Speculative Decoding for LLMs
di: Brown, Oscar, et al.
Pubblicazione: (2024)
di: Brown, Oscar, et al.
Pubblicazione: (2024)
PIS: Linking Importance Sampling and Attention Mechanisms for Efficient Prompt Compression
di: Chen, Lizhe, et al.
Pubblicazione: (2025)
di: Chen, Lizhe, et al.
Pubblicazione: (2025)
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference
di: Warner, Benjamin, et al.
Pubblicazione: (2024)
di: Warner, Benjamin, et al.
Pubblicazione: (2024)
Self-Supervised Prompt Optimization
di: Xiang, Jinyu, et al.
Pubblicazione: (2025)
di: Xiang, Jinyu, et al.
Pubblicazione: (2025)
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
di: Gim, In, et al.
Pubblicazione: (2023)
di: Gim, In, et al.
Pubblicazione: (2023)
Can Language Models Take A Hint? Prompting for Controllable Contextualized Commonsense Inference
di: Colon-Hernandez, Pedro, et al.
Pubblicazione: (2024)
di: Colon-Hernandez, Pedro, et al.
Pubblicazione: (2024)
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree
di: Gao, Xiangxiang, et al.
Pubblicazione: (2024)
di: Gao, Xiangxiang, et al.
Pubblicazione: (2024)
SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation
di: Orme, Michael, et al.
Pubblicazione: (2026)
di: Orme, Michael, et al.
Pubblicazione: (2026)
Fine-Tuned Language Models for Domain-Specific Summarization and Tagging
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
In-context Autoencoder for Context Compression in a Large Language Model
di: Ge, Tao, et al.
Pubblicazione: (2023)
di: Ge, Tao, et al.
Pubblicazione: (2023)
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
di: Huang, Sterling, et al.
Pubblicazione: (2026)
di: Huang, Sterling, et al.
Pubblicazione: (2026)
Structsum Generation for Faster Text Comprehension
di: Jain, Parag, et al.
Pubblicazione: (2024)
di: Jain, Parag, et al.
Pubblicazione: (2024)
HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
di: Shi, Zhiyuan, et al.
Pubblicazione: (2026)
Depth Registers Unlock W4A4 on SwiGLU: A Reader/Generator Decomposition
di: Liu, Ziyang
Pubblicazione: (2026)
di: Liu, Ziyang
Pubblicazione: (2026)
BayesPrompt: Prompting Large-Scale Pre-Trained Language Models on Few-shot Inference via Debiased Domain Abstraction
di: Li, Jiangmeng, et al.
Pubblicazione: (2024)
di: Li, Jiangmeng, et al.
Pubblicazione: (2024)
Inference-Aware Prompt Optimization for Aligning Black-Box Large Language Models
di: Mahmud, Saaduddin, et al.
Pubblicazione: (2025)
di: Mahmud, Saaduddin, et al.
Pubblicazione: (2025)
GRL-Prompt: Towards Knowledge Graph based Prompt Optimization via Reinforcement Learning
di: Liu, Yuze, et al.
Pubblicazione: (2024)
di: Liu, Yuze, et al.
Pubblicazione: (2024)
AhaKV: Adaptive Holistic Attention-Driven KV Cache Eviction for Efficient Inference of Large Language Models
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
di: Gu, Yifeng, et al.
Pubblicazione: (2025)
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
di: Ma, Guangyuan, et al.
Pubblicazione: (2025)
di: Ma, Guangyuan, et al.
Pubblicazione: (2025)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
di: Tsai, Yu-Che, et al.
Pubblicazione: (2026)
Revisiting In-context Learning Inference Circuit in Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2024)
di: Cho, Hakaze, et al.
Pubblicazione: (2024)
Communication Compression for Tensor Parallel LLM Inference
di: Hansen-Palmus, Jan, et al.
Pubblicazione: (2024)
di: Hansen-Palmus, Jan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
di: Kummer, Cornelius, et al.
Pubblicazione: (2026) -
Dynamic Compressing Prompts for Efficient Inference of Large Language Models
di: Hu, Jinwu, et al.
Pubblicazione: (2025) -
EFPC: Towards Efficient and Flexible Prompt Compression
di: Cao, Yun-Hao, et al.
Pubblicazione: (2025) -
Informed Routing in LLMs: Smarter Token-Level Computation for Faster Inference
di: Han, Chao, et al.
Pubblicazione: (2025) -
Discrete Prompt Compression with Reinforcement Learning
di: Jung, Hoyoun, et al.
Pubblicazione: (2023)