Prompt Compression in the Wild: Measuring Latency, Rate Adherence, and Quality for Faster LLM Inference
Fuente:
arXiv
Salvato in:
| Autori principali: | Kummer, Cornelius, Jurkschat, Lena, Färber, Michael, Vahdati, Sahar |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
di: Ma, Guangyuan, et al.
Pubblicazione: (2025)
di: Ma, Guangyuan, et al.
Pubblicazione: (2025)
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
di: Rogulsky, Steven, et al.
Pubblicazione: (2024)
di: Rogulsky, Steven, et al.
Pubblicazione: (2024)
Paths to Causality: Finding Informative Subgraphs Within Knowledge Graphs for Knowledge-Based Causal Discovery
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
di: Susanti, Yuni, et al.
Pubblicazione: (2025)
Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
di: Lepagnol, Pierre, et al.
Pubblicazione: (2025)
di: Lepagnol, Pierre, et al.
Pubblicazione: (2025)
LLM-Rec: Personalized Recommendation via Prompting Large Language Models
di: Lyu, Hanjia, et al.
Pubblicazione: (2023)
di: Lyu, Hanjia, et al.
Pubblicazione: (2023)
GenQREnsemble: Zero-Shot LLM Ensemble Prompting for Generative Query Reformulation
di: Dhole, Kaustubh, et al.
Pubblicazione: (2024)
di: Dhole, Kaustubh, et al.
Pubblicazione: (2024)
Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models
di: Kelber, Florian, et al.
Pubblicazione: (2026)
di: Kelber, Florian, et al.
Pubblicazione: (2026)
RDF-Based Structured Quality Assessment Representation of Multilingual LLM Evaluations
di: Gwozdz, Jonas, et al.
Pubblicazione: (2025)
di: Gwozdz, Jonas, et al.
Pubblicazione: (2025)
No Free Lunch in Active Learning: LLM Embedding Quality Dictates Query Strategy Success
di: Rauch, Lukas, et al.
Pubblicazione: (2025)
di: Rauch, Lukas, et al.
Pubblicazione: (2025)
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
di: Xu, Ziyang, et al.
Pubblicazione: (2024)
OSCAR: Online Soft Compression And Reranking
di: Louis, Maxime, et al.
Pubblicazione: (2025)
di: Louis, Maxime, et al.
Pubblicazione: (2025)
Benchmarking Prompt Sensitivity in Large Language Models
di: Razavi, Amirhossein, et al.
Pubblicazione: (2025)
di: Razavi, Amirhossein, et al.
Pubblicazione: (2025)
Spectral Tempering for Embedding Compression in Dense Passage Retrieval
di: Li, Yongkang, et al.
Pubblicazione: (2026)
di: Li, Yongkang, et al.
Pubblicazione: (2026)
PISCO: Pretty Simple Compression for Retrieval-Augmented Generation
di: Louis, Maxime, et al.
Pubblicazione: (2025)
di: Louis, Maxime, et al.
Pubblicazione: (2025)
Modular Representation Compression: Adapting LLMs for Efficient and Effective Recommendations
di: Xi, Yunjia, et al.
Pubblicazione: (2026)
di: Xi, Yunjia, et al.
Pubblicazione: (2026)
SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
di: Jin, Yiqiao, et al.
Pubblicazione: (2025)
di: Jin, Yiqiao, et al.
Pubblicazione: (2025)
ECoRAG: Evidentiality-guided Compression for Long Context RAG
di: Jeong, Yeonseok, et al.
Pubblicazione: (2025)
di: Jeong, Yeonseok, et al.
Pubblicazione: (2025)
Chained Prompting for Better Systematic Review Search Strategies
di: Nasser, Fatima, et al.
Pubblicazione: (2025)
di: Nasser, Fatima, et al.
Pubblicazione: (2025)
EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation
di: Hwang, Taeho, et al.
Pubblicazione: (2024)
di: Hwang, Taeho, et al.
Pubblicazione: (2024)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
di: Thellmann, Klaudia, et al.
Pubblicazione: (2026)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2026)
xRAG: Extreme Context Compression for Retrieval-augmented Generation with One Token
di: Cheng, Xin, et al.
Pubblicazione: (2024)
di: Cheng, Xin, et al.
Pubblicazione: (2024)
Better by Comparison: Retrieval-Augmented Contrastive Reasoning for Automatic Prompt Optimization
di: Lee, Juhyeon, et al.
Pubblicazione: (2025)
di: Lee, Juhyeon, et al.
Pubblicazione: (2025)
UNH at CheckThat! 2025: Fine-tuning Vs Prompting in Claim Extraction
di: Wilder, Joe, et al.
Pubblicazione: (2025)
di: Wilder, Joe, et al.
Pubblicazione: (2025)
The Synergy of Automated Pipelines with Prompt Engineering and Generative AI in Web Crawling
di: Huang, Chau-Jian
Pubblicazione: (2024)
di: Huang, Chau-Jian
Pubblicazione: (2024)
Can Artificial Intelligence Generate Quality Research Topics Reflecting Patient Concerns?
di: Kim, Jiyeong, et al.
Pubblicazione: (2024)
di: Kim, Jiyeong, et al.
Pubblicazione: (2024)
Can we Retrieve Everything All at Once? ARM: An Alignment-Oriented LLM-based Retrieval Method
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
di: Chen, Peter Baile, et al.
Pubblicazione: (2025)
StealthRank: LLM Ranking Manipulation via Stealthy Prompt Optimization
di: Tang, Yiming, et al.
Pubblicazione: (2025)
di: Tang, Yiming, et al.
Pubblicazione: (2025)
When "Better" Prompts Hurt: Evaluation-Driven Iteration for LLM Applications
di: Commey, Daniel
Pubblicazione: (2026)
di: Commey, Daniel
Pubblicazione: (2026)
LLM-RankFusion: Mitigating Intrinsic Inconsistency in LLM-based Ranking
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
di: Zeng, Yifan, et al.
Pubblicazione: (2024)
Rethinking Soft Compression in Retrieval-Augmented Generation: A Query-Conditioned Selector Perspective
di: Liu, Yunhao, et al.
Pubblicazione: (2026)
di: Liu, Yunhao, et al.
Pubblicazione: (2026)
PERSOMA: PERsonalized SOft ProMpt Adapter Architecture for Personalized Language Prompting
di: Hebert, Liam, et al.
Pubblicazione: (2024)
di: Hebert, Liam, et al.
Pubblicazione: (2024)
Knowledge Compression via Question Generation: Enhancing Multihop Document Retrieval without Fine-tuning
di: Eponon, Anvi Alex, et al.
Pubblicazione: (2025)
di: Eponon, Anvi Alex, et al.
Pubblicazione: (2025)
PromptLink: Leveraging Large Language Models for Cross-Source Biomedical Concept Linking
di: Xie, Yuzhang, et al.
Pubblicazione: (2024)
di: Xie, Yuzhang, et al.
Pubblicazione: (2024)
Horizon Scans can be accelerated using novel information retrieval and artificial intelligence tools
di: Schmidt, Lena, et al.
Pubblicazione: (2025)
di: Schmidt, Lena, et al.
Pubblicazione: (2025)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
RecGPT: Generative Personalized Prompts for Sequential Recommendation via ChatGPT Training Paradigm
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
di: Zhang, Yabin, et al.
Pubblicazione: (2024)
Little Giants: Synthesizing High-Quality Embedding Data at Scale
di: Chen, Haonan, et al.
Pubblicazione: (2024)
di: Chen, Haonan, et al.
Pubblicazione: (2024)
LLM2IR: simple unsupervised contrastive learning makes long-context LLM great retriever
di: Yang, Xiaocong
Pubblicazione: (2025)
di: Yang, Xiaocong
Pubblicazione: (2025)
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
di: Wang, Guangzhi, et al.
Pubblicazione: (2025)
di: Wang, Guangzhi, et al.
Pubblicazione: (2025)
Enhanced Facet Generation with LLM Editing
di: Lee, Joosung, et al.
Pubblicazione: (2024)
di: Lee, Joosung, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LightRetriever: A LLM-based Text Retrieval Architecture with Extremely Faster Query Inference
di: Ma, Guangyuan, et al.
Pubblicazione: (2025) -
The Effects of Hallucinations in Synthetic Training Data for Relation Extraction
di: Rogulsky, Steven, et al.
Pubblicazione: (2024) -
Paths to Causality: Finding Informative Subgraphs Within Knowledge Graphs for Knowledge-Based Causal Discovery
di: Susanti, Yuni, et al.
Pubblicazione: (2025) -
Leveraging Information Retrieval to Enhance Spoken Language Understanding Prompts in Few-Shot Learning
di: Lepagnol, Pierre, et al.
Pubblicazione: (2025) -
LLM-Rec: Personalized Recommendation via Prompting Large Language Models
di: Lyu, Hanjia, et al.
Pubblicazione: (2023)