The Compression Paradox in LLM Inference: Provider-Dependent Energy Effects of Prompt Compression
Fuente:
arXiv
Salvato in:
| Autore principale: | Johnson, Warren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
di: Johnson, Warren, et al.
Pubblicazione: (2026)
di: Johnson, Warren, et al.
Pubblicazione: (2026)
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
di: Yankun, Hong, et al.
Pubblicazione: (2025)
di: Yankun, Hong, et al.
Pubblicazione: (2025)
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
di: Park, Seungcheol, et al.
Pubblicazione: (2025)
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)
Visual Word Sense Disambiguation with CLIP through Dual-Channel Text Prompting and Image Augmentations
di: Bhattacharya, Shamik, et al.
Pubblicazione: (2026)
di: Bhattacharya, Shamik, et al.
Pubblicazione: (2026)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
di: Chen, Jiaju, et al.
Pubblicazione: (2025)
di: Chen, Jiaju, et al.
Pubblicazione: (2025)
A Comprehensive Survey of Compression Algorithms for Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2024)
di: Park, Seungcheol, et al.
Pubblicazione: (2024)
LLM-Assisted Crisis Management: Building Advanced LLM Platforms for Effective Emergency Response and Public Collaboration
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
di: Otal, Hakan T., et al.
Pubblicazione: (2024)
Understanding and Improving Information Preservation in Prompt Compression for LLMs
di: Łajewska, Weronika, et al.
Pubblicazione: (2025)
di: Łajewska, Weronika, et al.
Pubblicazione: (2025)
The Paradox of Poetic Intent in Back-Translation: Evaluating the Quality of Large Language Models in Chinese Translation
di: Weigang, Li, et al.
Pubblicazione: (2025)
di: Weigang, Li, et al.
Pubblicazione: (2025)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
di: Li, Mingda, et al.
Pubblicazione: (2024)
di: Li, Mingda, et al.
Pubblicazione: (2024)
A Primer on Large Language Models and their Limitations
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
di: Johnson, Sandra, et al.
Pubblicazione: (2024)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
di: Vosoughi, Ali, et al.
Pubblicazione: (2025)
Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
di: Arabov, Mullosharaf K.
Pubblicazione: (2026)
LangMARL: Natural Language Multi-Agent Reinforcement Learning
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
di: Yao, Huaiyuan, et al.
Pubblicazione: (2026)
LLMs as Deceptive Agents: How Role-Based Prompting Induces Semantic Ambiguity in Puzzle Tasks
di: Yoo, Seunghyun
Pubblicazione: (2025)
di: Yoo, Seunghyun
Pubblicazione: (2025)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
Piloting Copilot, Codex, and StarCoder2: Hot Temperature, Cold Prompts, or Black Magic?
di: Döderlein, Jean-Baptiste, et al.
Pubblicazione: (2022)
di: Döderlein, Jean-Baptiste, et al.
Pubblicazione: (2022)
BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation
di: Li, Xuan, et al.
Pubblicazione: (2026)
di: Li, Xuan, et al.
Pubblicazione: (2026)
Multicultural Spyfall: Assessing LLMs through Dynamic Multilingual Social Deduction Game
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2026)
The Privileged Students: On the Value of Initialization in Multilingual Knowledge Distillation
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2024)
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
di: Andrylie, Lyzander Marciano, et al.
Pubblicazione: (2025)
Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
di: Land, Sander, et al.
Pubblicazione: (2024)
di: Land, Sander, et al.
Pubblicazione: (2024)
Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art
di: Ikoma, Hayato, et al.
Pubblicazione: (2025)
di: Ikoma, Hayato, et al.
Pubblicazione: (2025)
COPAL-ID: Indonesian Language Reasoning with Local Culture and Nuances
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2023)
Knowledge Editing for Large Language Model with Knowledge Neuronal Ensemble
di: Li, Yongchang, et al.
Pubblicazione: (2024)
di: Li, Yongchang, et al.
Pubblicazione: (2024)
A Survey on Hypothesis Generation for Scientific Discovery in the Era of Large Language Models
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
di: Alkan, Atilla Kaan, et al.
Pubblicazione: (2025)
Which Pieces Does Unigram Tokenization Really Need?
di: Land, Sander, et al.
Pubblicazione: (2025)
di: Land, Sander, et al.
Pubblicazione: (2025)
Duluth at SemEval-2025 Task 7: TF-IDF with Optimized Vector Dimensions for Multilingual Fact-Checked Claim Retrieval
di: Syed, Shujauddin, et al.
Pubblicazione: (2025)
di: Syed, Shujauddin, et al.
Pubblicazione: (2025)
Knesset-DictaBERT: A Hebrew Language Model for Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
IteRABRe: Iterative Recovery-Aided Block Reduction
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
di: Wibowo, Haryo Akbarianto, et al.
Pubblicazione: (2025)
Number Representations in LLMs: A Computational Parallel to Human Perception
di: AlquBoj, H. V., et al.
Pubblicazione: (2025)
di: AlquBoj, H. V., et al.
Pubblicazione: (2025)
Healthy LLMs? Benchmarking LLM Knowledge of UK Government Public Health Information
di: Harris, Joshua, et al.
Pubblicazione: (2025)
di: Harris, Joshua, et al.
Pubblicazione: (2025)
Math Natural Language Inference: this should be easy!
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
Separate Before You Compress: The WWHO Tokenization Architecture
di: Darshana, Kusal
Pubblicazione: (2026)
di: Darshana, Kusal
Pubblicazione: (2026)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
di: Liu, Aiwei, et al.
Pubblicazione: (2024)
From Scarcity to Efficiency: Investigating the Effects of Data Augmentation on African Machine Translation
di: Oduwole, Mardiyyah, et al.
Pubblicazione: (2025)
di: Oduwole, Mardiyyah, et al.
Pubblicazione: (2025)
Surfing the modeling of PoS taggers in low-resource scenarios
di: Ferro, Manuel Vilares, et al.
Pubblicazione: (2024)
di: Ferro, Manuel Vilares, et al.
Pubblicazione: (2024)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
di: Johnson, Warren
Pubblicazione: (2026) -
Prompt Compression in Production Task Orchestration: A Pre-Registered Randomized Trial
di: Johnson, Warren, et al.
Pubblicazione: (2026) -
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
di: Yankun, Hong, et al.
Pubblicazione: (2025) -
Unifying Uniform and Binary-coding Quantization for Accurate Compression of Large Language Models
di: Park, Seungcheol, et al.
Pubblicazione: (2025) -
Tokenization Is More Than Compression
di: Schmidt, Craig W., et al.
Pubblicazione: (2024)