EvoP: Robust LLM Inference via Evolutionary Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Shangyu, Du, Hongchao, Xiong, Ying, Chen, Shuai, Kuo, Tei-Wei, Guan, Nan, Xue, Chun Jason |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024)
by: Huang, Lianming, et al.
Published: (2024)
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025)
by: Du, Hongchao, et al.
Published: (2025)
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
by: Chen, Yixin, et al.
Published: (2026)
by: Chen, Yixin, et al.
Published: (2026)
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
Beyond Semantic Similarity: Reducing Unnecessary API Calls via Behavior-Aligned Retriever
by: Chen, Yixin, et al.
Published: (2025)
by: Chen, Yixin, et al.
Published: (2025)
CHESS: Optimizing LLM Inference via Channel-Wise Thresholding and Selective Sparsification
by: He, Junhui, et al.
Published: (2024)
by: He, Junhui, et al.
Published: (2024)
On the Compressibility of Quantized Large Language Models
by: Mao, Yu, et al.
Published: (2024)
by: Mao, Yu, et al.
Published: (2024)
When Compression Meets Model Compression: Memory-Efficient Double Compression for Large Language Models
by: Wang, Weilan, et al.
Published: (2025)
by: Wang, Weilan, et al.
Published: (2025)
A$^2$ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
by: He, Junhui, et al.
Published: (2025)
by: He, Junhui, et al.
Published: (2025)
EvoMU: Evolutionary Machine Unlearning
by: Batorski, Pawel, et al.
Published: (2026)
by: Batorski, Pawel, et al.
Published: (2026)
CoT-Evo: Evolutionary Distillation of Chain-of-Thought for Scientific Reasoning
by: Feng, Kehua, et al.
Published: (2025)
by: Feng, Kehua, et al.
Published: (2025)
KCoEvo: A Knowledge Graph Augmented Framework for Evolutionary Code Generation
by: Kang, Jiazhen, et al.
Published: (2026)
by: Kang, Jiazhen, et al.
Published: (2026)
Easz: An Agile Transformer-based Image Compression Framework for Resource-constrained IoTs
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
Pruning as a Domain-specific LLM Extractor
by: Zhang, Nan, et al.
Published: (2024)
by: Zhang, Nan, et al.
Published: (2024)
ClawMobile: Rethinking Smartphone-Native Agentic Systems
by: Du, Hongchao, et al.
Published: (2026)
by: Du, Hongchao, et al.
Published: (2026)
Prune&Comp: Free Lunch for Layer-Pruned LLMs via Iterative Pruning with Magnitude Compensation
by: Chen, Xinrui, et al.
Published: (2025)
by: Chen, Xinrui, et al.
Published: (2025)
SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token Pruning
by: Long, Lingkun, et al.
Published: (2025)
by: Long, Lingkun, et al.
Published: (2025)
EvoPool: Evolutionary Programmatic Annotation for Label-Efficient Specialized Supervision
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation
by: Li, Ting-Wei, et al.
Published: (2026)
by: Li, Ting-Wei, et al.
Published: (2026)
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Beyond One-Size-Fits-All Pruning via Evolutionary Metric Search for Large Language Models
by: Liu, Shuqi, et al.
Published: (2025)
by: Liu, Shuqi, et al.
Published: (2025)
EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
by: Guo, Qingyan, et al.
Published: (2023)
by: Guo, Qingyan, et al.
Published: (2023)
EvoWiki: Evaluating LLMs on Evolving Knowledge
by: Tang, Wei, et al.
Published: (2024)
by: Tang, Wei, et al.
Published: (2024)
EvoRoute: Experience-Driven Self-Routing LLM Agent Systems
by: Zhang, Guibin, et al.
Published: (2026)
by: Zhang, Guibin, et al.
Published: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
by: Fu, Qichen, et al.
Published: (2024)
by: Fu, Qichen, et al.
Published: (2024)
AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data
by: Wu, JiaRu, et al.
Published: (2025)
by: Wu, JiaRu, et al.
Published: (2025)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
by: Li, Jianwei, et al.
Published: (2023)
by: Li, Jianwei, et al.
Published: (2023)
ContiguousKV: Accelerating LLM Prefill with Granularity-Aligned KV Cache Management
by: Zou, Jing, et al.
Published: (2026)
by: Zou, Jing, et al.
Published: (2026)
EvoGPT-f: An Evolutionary GPT Framework for Benchmarking Formal Math Languages
by: Mercer, Johnathan
Published: (2024)
by: Mercer, Johnathan
Published: (2024)
Efficient LLM Inference using Dynamic Input Pruning and Cache-Aware Masking
by: Federici, Marco, et al.
Published: (2024)
by: Federici, Marco, et al.
Published: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
by: Yu, Bohan, et al.
Published: (2025)
by: Yu, Bohan, et al.
Published: (2025)
Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM
by: Xia, Chunqiu Steven, et al.
Published: (2024)
by: Xia, Chunqiu Steven, et al.
Published: (2024)
Lossless Compression of Large Language Model-Generated Text via Next-Token Prediction
by: Mao, Yu, et al.
Published: (2025)
by: Mao, Yu, et al.
Published: (2025)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
by: Taniguchi, Rei, et al.
Published: (2026)
by: Taniguchi, Rei, et al.
Published: (2026)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
EvoDiagram: Agentic Editable Diagram Creation via Design Expertise Evolution
by: Wang, Tianfu, et al.
Published: (2026)
by: Wang, Tianfu, et al.
Published: (2026)
Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
by: Hu, Haibo, et al.
Published: (2025)
by: Hu, Haibo, et al.
Published: (2025)
Similar Items
-
RAEE: A Robust Retrieval-Augmented Early Exit Framework for Efficient Inference
by: Huang, Lianming, et al.
Published: (2024) -
FlexInfer: Breaking Memory Constraint via Flexible and Efficient Offloading for On-Device LLM Inference
by: Du, Hongchao, et al.
Published: (2025) -
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
by: Chen, Yixin, et al.
Published: (2026) -
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024) -
Retrieval-Augmented Generation for Natural Language Processing: A Survey
by: Wu, Shangyu, et al.
Published: (2024)