The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Binqi, Jin, Lier, Cai, Hanyu, Hu, Lan, Xin, Yuting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
von: Cai, Hanyu, et al.
Veröffentlicht: (2025)
von: Cai, Hanyu, et al.
Veröffentlicht: (2025)
Unified Context Evolution for LLM Agents
von: Zhu, Zixuan, et al.
Veröffentlicht: (2026)
von: Zhu, Zixuan, et al.
Veröffentlicht: (2026)
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
von: Xu, Haolei, et al.
Veröffentlicht: (2025)
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
von: Xu, Shiyi, et al.
Veröffentlicht: (2025)
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
von: Ye, Yongshi, et al.
Veröffentlicht: (2026)
von: Ye, Yongshi, et al.
Veröffentlicht: (2026)
Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling
von: Fashi, Parsa Ashrafi, et al.
Veröffentlicht: (2026)
von: Fashi, Parsa Ashrafi, et al.
Veröffentlicht: (2026)
Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
von: Xu, Haoran, et al.
Veröffentlicht: (2024)
Single LLM, Multiple Roles: A Unified Retrieval-Augmented Generation Framework Using Role-Specific Token Optimization
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
von: Zhu, Yutao, et al.
Veröffentlicht: (2025)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
von: Jing, Huihao, et al.
Veröffentlicht: (2026)
Path-LLM: A Shortest-Path-based LLM Learning for Unified Graph Representation
von: Shang, Wenbo, et al.
Veröffentlicht: (2024)
von: Shang, Wenbo, et al.
Veröffentlicht: (2024)
promptolution: A Unified, Modular Framework for Prompt Optimization
von: Zehle, Tom, et al.
Veröffentlicht: (2025)
von: Zehle, Tom, et al.
Veröffentlicht: (2025)
LLM Hallucination Detection: HSAD
von: Li, JinXin, et al.
Veröffentlicht: (2025)
von: Li, JinXin, et al.
Veröffentlicht: (2025)
AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2025)
von: Chen, Xuanzhong, et al.
Veröffentlicht: (2025)
Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agents
von: Pollertlam, Natchanon, et al.
Veröffentlicht: (2026)
von: Pollertlam, Natchanon, et al.
Veröffentlicht: (2026)
Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
von: Sakhawat, Adib, et al.
Veröffentlicht: (2026)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
von: Yao, Dingyu, et al.
Veröffentlicht: (2025)
Democratizing LLM Efficiency: From Hyperscale Optimizations to Universal Deployability
von: Huang, Hen-Hsen
Veröffentlicht: (2025)
von: Huang, Hen-Hsen
Veröffentlicht: (2025)
MemBoost: A Memory-Boosted Framework for Cost-Aware LLM Inference
von: Köster, Joris, et al.
Veröffentlicht: (2026)
von: Köster, Joris, et al.
Veröffentlicht: (2026)
Dissecting the Impact of Mobile DVFS Governors on LLM Inference Performance and Energy Efficiency
von: Zhang, Zongpu, et al.
Veröffentlicht: (2025)
von: Zhang, Zongpu, et al.
Veröffentlicht: (2025)
MVSS: A Unified Framework for Multi-View Structured Survey Generation
von: Liu, Yinqi, et al.
Veröffentlicht: (2026)
von: Liu, Yinqi, et al.
Veröffentlicht: (2026)
Beyond GPT-5: Making LLMs Cheaper and Better via Performance-Efficiency Optimized Routing
von: Zhang, Yiqun, et al.
Veröffentlicht: (2025)
von: Zhang, Yiqun, et al.
Veröffentlicht: (2025)
Less Context, Same Performance: A RAG Framework for Resource-Efficient LLM-Based Clinical NLP
von: Cheetirala, Satya Narayana, et al.
Veröffentlicht: (2025)
von: Cheetirala, Satya Narayana, et al.
Veröffentlicht: (2025)
La RoSA: Enhancing LLM Efficiency via Layerwise Rotated Sparse Activation
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
Learning from Fine-Grained Visual Discrepancies: Mitigating Multimodal Hallucinations via In-Context Visual Contrastive Optimization
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
von: Deng, Haolin, et al.
Veröffentlicht: (2026)
Hierarchical Chain-of-Thought Prompting: Enhancing LLM Reasoning Performance and Efficiency
von: Huang, Xingshuai, et al.
Veröffentlicht: (2026)
von: Huang, Xingshuai, et al.
Veröffentlicht: (2026)
Towards Optimizing the Costs of LLM Usage
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
von: Shekhar, Shivanshu, et al.
Veröffentlicht: (2024)
Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
von: Shi, Luohe, et al.
Veröffentlicht: (2024)
Do LLMs Know about Hallucination? An Empirical Investigation of LLM's Hidden States
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
von: Duan, Hanyu, et al.
Veröffentlicht: (2024)
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report
von: Lab, Shanghai AI, et al.
Veröffentlicht: (2025)
von: Lab, Shanghai AI, et al.
Veröffentlicht: (2025)
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
von: Matta, Shiho, et al.
Veröffentlicht: (2024)
von: Matta, Shiho, et al.
Veröffentlicht: (2024)
Reward-aware Preference Optimization: A Unified Mathematical Framework for Model Alignment
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
von: Sun, Shengyang, et al.
Veröffentlicht: (2025)
DEPO: Dual-Efficiency Preference Optimization for LLM Agents
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
U-NIAH: Unified RAG and LLM Evaluation for Long Context Needle-In-A-Haystack
von: Gao, Yunfan, et al.
Veröffentlicht: (2025)
von: Gao, Yunfan, et al.
Veröffentlicht: (2025)
Compress the Context, Keep the Commitments: A Formal Framework for Verifiable LLM Context Compression
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
von: Trukhina, Natalia, et al.
Veröffentlicht: (2026)
AutoContext: Instance-Level Context Learning for LLM Agents
von: Cai, Kuntai, et al.
Veröffentlicht: (2025)
von: Cai, Kuntai, et al.
Veröffentlicht: (2025)
Aioli: A Unified Optimization Framework for Language Model Data Mixing
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
von: Chen, Mayee F., et al.
Veröffentlicht: (2024)
Context Discipline and Performance Correlation: Analyzing LLM Performance and Quality Degradation Under Varying Context Lengths
von: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Veröffentlicht: (2025)
von: Ponnusamy, Ahilan Ayyachamy Nadar, et al.
Veröffentlicht: (2025)
SVD-LLM V2: Optimizing Singular Value Truncation for Large Language Model Compression
von: Wang, Xin, et al.
Veröffentlicht: (2025)
von: Wang, Xin, et al.
Veröffentlicht: (2025)
GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistant
von: Shen, Zhuokang, et al.
Veröffentlicht: (2026)
von: Shen, Zhuokang, et al.
Veröffentlicht: (2026)
UniCBE: An Uniformity-driven Comparing Based Evaluation Framework with Unified Multi-Objective Optimization
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMA
von: Cai, Hanyu, et al.
Veröffentlicht: (2025) -
Unified Context Evolution for LLM Agents
von: Zhu, Zixuan, et al.
Veröffentlicht: (2026) -
EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering
von: Xu, Haolei, et al.
Veröffentlicht: (2025) -
ICPC-Eval: Probing the Frontiers of LLM Reasoning with Competitive Programming Contests
von: Xu, Shiyi, et al.
Veröffentlicht: (2025) -
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory
von: Ye, Yongshi, et al.
Veröffentlicht: (2026)