RAP: Runtime Adaptive Pruning for LLM Inference
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liu, Huanrong, Tian, Chunlin, Wei, Xuyang, Li, Qingbiao, Li, Li |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Less is More: Resource-Efficient Low-Rank Adaptation
par: Tian, Chunlin, et autres
Publié: (2025)
par: Tian, Chunlin, et autres
Publié: (2025)
RAP: KV-Cache Compression via RoPE-Aligned Pruning
par: Xin, Jihao, et autres
Publié: (2026)
par: Xin, Jihao, et autres
Publié: (2026)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
par: Li, Junjie, et autres
Publié: (2026)
par: Li, Junjie, et autres
Publié: (2026)
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
par: Taniguchi, Rei, et autres
Publié: (2026)
par: Taniguchi, Rei, et autres
Publié: (2026)
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
par: Li, Yingru, et autres
Publié: (2025)
par: Li, Yingru, et autres
Publié: (2025)
AFD-STA: Adaptive Filtering Denoising with Spatiotemporal Attention for Chaotic System Prediction
par: Gong, Chunlin, et autres
Publié: (2025)
par: Gong, Chunlin, et autres
Publié: (2025)
Heterogeneity-Aware Coordination for Federated Learning via Stitching Pre-trained blocks
par: Zhan, Shichen, et autres
Publié: (2024)
par: Zhan, Shichen, et autres
Publié: (2024)
Martingale Foresight Sampling: A Principled Approach to Inference-Time LLM Decoding
par: Li, Huayu, et autres
Publié: (2026)
par: Li, Huayu, et autres
Publié: (2026)
PruneSymNet: A Symbolic Neural Network and Pruning Algorithm for Symbolic Regression
par: Wu, Min, et autres
Publié: (2024)
par: Wu, Min, et autres
Publié: (2024)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
par: Li, Tenghui, et autres
Publié: (2025)
par: Li, Tenghui, et autres
Publié: (2025)
Exploring Federated Pruning for Large Language Models
par: Guo, Pengxin, et autres
Publié: (2025)
par: Guo, Pengxin, et autres
Publié: (2025)
TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI
par: Oh, Hyunwoo, et autres
Publié: (2026)
par: Oh, Hyunwoo, et autres
Publié: (2026)
MASteer: Multi-Agent Adaptive Steer Strategy for End-to-End LLM Trustworthiness Repair
par: Li, Changqing, et autres
Publié: (2025)
par: Li, Changqing, et autres
Publié: (2025)
RAP: Retrieval-Augmented Planning with Contextual Memory for Multimodal LLM Agents
par: Kagaya, Tomoyuki, et autres
Publié: (2024)
par: Kagaya, Tomoyuki, et autres
Publié: (2024)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
par: Noh, Kangjun, et autres
Publié: (2026)
par: Noh, Kangjun, et autres
Publié: (2026)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
par: Fu, Qichen, et autres
Publié: (2024)
par: Fu, Qichen, et autres
Publié: (2024)
Heterogeneous Graph Prompt Learning via Adaptive Weight Pruning
par: Wei, Chu-Yuan, et autres
Publié: (2025)
par: Wei, Chu-Yuan, et autres
Publié: (2025)
Adaptive Layer Splitting for Wireless LLM Inference in Edge Computing: A Model-Based Reinforcement Learning Approach
par: Chen, Yuxuan, et autres
Publié: (2024)
par: Chen, Yuxuan, et autres
Publié: (2024)
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
par: Li, Guangyan, et autres
Publié: (2024)
par: Li, Guangyan, et autres
Publié: (2024)
Prompt-prompted Adaptive Structured Pruning for Efficient LLM Generation
par: Dong, Harry, et autres
Publié: (2024)
par: Dong, Harry, et autres
Publié: (2024)
EMP: Enhance Memory in Data Pruning
par: Xiao, Jinying, et autres
Publié: (2024)
par: Xiao, Jinying, et autres
Publié: (2024)
Exploiting Adaptive Channel Pruning for Communication-Efficient Split Learning
par: Tan, Jialei, et autres
Publié: (2026)
par: Tan, Jialei, et autres
Publié: (2026)
Sharpening the Spear: Adaptive Expert-Guided Adversarial Attack Against DRL-based Autonomous Driving Policies
par: Fan, Junchao, et autres
Publié: (2025)
par: Fan, Junchao, et autres
Publié: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
par: Qin, Jiayu, et autres
Publié: (2025)
par: Qin, Jiayu, et autres
Publié: (2025)
UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs
par: Neth, Ashe, et autres
Publié: (2025)
par: Neth, Ashe, et autres
Publié: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
par: Yuan, Xin, et autres
Publié: (2025)
par: Yuan, Xin, et autres
Publié: (2025)
Rotated Runtime Smooth: Training-Free Activation Smoother for accurate INT4 inference
par: Yi, Ke, et autres
Publié: (2024)
par: Yi, Ke, et autres
Publié: (2024)
DP-LLM: Runtime Model Adaptation with Dynamic Layer-wise Precision Assignment
par: Kwon, Sangwoo, et autres
Publié: (2025)
par: Kwon, Sangwoo, et autres
Publié: (2025)
Hybrid Dynamic Pruning: A Pathway to Efficient Transformer Inference
par: Jaradat, Ghadeer, et autres
Publié: (2024)
par: Jaradat, Ghadeer, et autres
Publié: (2024)
Towards Mitigating Architecture Overfitting on Distilled Datasets
par: Zhong, Xuyang, et autres
Publié: (2023)
par: Zhong, Xuyang, et autres
Publié: (2023)
Runtime Tunable Tsetlin Machines for Edge Inference on eFPGAs
par: Rahman, Tousif, et autres
Publié: (2025)
par: Rahman, Tousif, et autres
Publié: (2025)
RaPA: Enhancing Transferable Targeted Attacks via Random Parameter Pruning
par: Su, Tongrui, et autres
Publié: (2025)
par: Su, Tongrui, et autres
Publié: (2025)
GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models
par: Liu, Xiaoyun, et autres
Publié: (2026)
par: Liu, Xiaoyun, et autres
Publié: (2026)
AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block Size
par: Lu, Guanxi, et autres
Publié: (2025)
par: Lu, Guanxi, et autres
Publié: (2025)
Uncovering Capabilities of Model Pruning in Graph Contrastive Learning
par: Wu, Junran, et autres
Publié: (2024)
par: Wu, Junran, et autres
Publié: (2024)
RAMP: Reinforcement Adaptive Mixed Precision Quantization for Efficient On Device LLM Inference
par: Gautam, Arpit Singh, et autres
Publié: (2026)
par: Gautam, Arpit Singh, et autres
Publié: (2026)
CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration
par: Han, Yuning, et autres
Publié: (2026)
par: Han, Yuning, et autres
Publié: (2026)
LASER: Language Model Regression for Semi-Structured Workflow Resource and Runtime Estimation
par: Yin, Yuxuan, et autres
Publié: (2025)
par: Yin, Yuxuan, et autres
Publié: (2025)
Scaling Capability in Token Space: An Analysis of Large Vision Language Model
par: Li, Tenghui, et autres
Publié: (2024)
par: Li, Tenghui, et autres
Publié: (2024)
Adaptive Pruning for Large Language Models with Structural Importance Awareness
par: Zheng, Haotian, et autres
Publié: (2024)
par: Zheng, Haotian, et autres
Publié: (2024)
Documents similaires
-
Less is More: Resource-Efficient Low-Rank Adaptation
par: Tian, Chunlin, et autres
Publié: (2025) -
RAP: KV-Cache Compression via RoPE-Aligned Pruning
par: Xin, Jihao, et autres
Publié: (2026) -
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
par: Li, Junjie, et autres
Publié: (2026) -
Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
par: Taniguchi, Rei, et autres
Publié: (2026) -
Dynamic Vocabulary Pruning: Stable LLM-RL by Taming the Tail
par: Li, Yingru, et autres
Publié: (2025)