ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Junjie, Lou, Jiong, Li, Jie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
por: Li, Miao, et al.
Publicado: (2026)
por: Li, Miao, et al.
Publicado: (2026)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
por: Zhu, Yu, et al.
Publicado: (2024)
por: Zhu, Yu, et al.
Publicado: (2024)
Predicting LLM Reasoning Performance with Small Proxy Model
por: Koh, Woosung, et al.
Publicado: (2025)
por: Koh, Woosung, et al.
Publicado: (2025)
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
por: Fu, Qichen, et al.
Publicado: (2024)
por: Fu, Qichen, et al.
Publicado: (2024)
Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
por: Liu, Zhihao, et al.
Publicado: (2025)
por: Liu, Zhihao, et al.
Publicado: (2025)
OptiProxy-NAS: Optimization Proxy based End-to-End Neural Architecture Search
por: Lyu, Bo, et al.
Publicado: (2025)
por: Lyu, Bo, et al.
Publicado: (2025)
KV Admission: Learning What to Write for Efficient Long-Context Inference
por: Huang, Yen-Chieh, et al.
Publicado: (2025)
por: Huang, Yen-Chieh, et al.
Publicado: (2025)
Feature Alignment: Rethinking Efficient Active Learning via Proxy in the Context of Pre-trained Models
por: Wen, Ziting, et al.
Publicado: (2024)
por: Wen, Ziting, et al.
Publicado: (2024)
SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM Inference
por: Zhao, Yi, et al.
Publicado: (2025)
por: Zhao, Yi, et al.
Publicado: (2025)
MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
por: Li, Kunxi, et al.
Publicado: (2025)
por: Li, Kunxi, et al.
Publicado: (2025)
TTKV: Temporal-Tiered KV Cache for Long-Context LLM Inference
por: Dzikanyanga, Gradwell, et al.
Publicado: (2026)
por: Dzikanyanga, Gradwell, et al.
Publicado: (2026)
FreeKV: Boosting KV Cache Retrieval for Efficient LLM Inference
por: Liu, Guangda, et al.
Publicado: (2025)
por: Liu, Guangda, et al.
Publicado: (2025)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
por: Li, Dongfang, et al.
Publicado: (2026)
por: Li, Dongfang, et al.
Publicado: (2026)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
por: Butler, Landon, et al.
Publicado: (2025)
por: Butler, Landon, et al.
Publicado: (2025)
FedProxy: Federated Fine-Tuning of LLMs via Proxy SLMs and Heterogeneity-Aware Fusion
por: Fan, Tao, et al.
Publicado: (2026)
por: Fan, Tao, et al.
Publicado: (2026)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
por: Wang, Guangtao, et al.
Publicado: (2025)
por: Wang, Guangtao, et al.
Publicado: (2025)
High Rank Matrix Completion via Grassmannian Proxy Fusion
por: Li, Huanran, et al.
Publicado: (2026)
por: Li, Huanran, et al.
Publicado: (2026)
IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference
por: Liu, Guanming, et al.
Publicado: (2026)
por: Liu, Guanming, et al.
Publicado: (2026)
RAP: Runtime Adaptive Pruning for LLM Inference
por: Liu, Huanrong, et al.
Publicado: (2025)
por: Liu, Huanrong, et al.
Publicado: (2025)
E-TCAV: Formalizing Penultimate Proxies for Efficient Concept Based Interpretability
por: Aslam, Hasib, et al.
Publicado: (2026)
por: Aslam, Hasib, et al.
Publicado: (2026)
LouisKV: Efficient KV Cache Retrieval for Long Input-Output Sequences
por: Wu, Wenbo, et al.
Publicado: (2025)
por: Wu, Wenbo, et al.
Publicado: (2025)
OEP: Poisoning Self-Evolving LLM Agents via Locally Correct but Non-Transferable Experiences
por: Wang, Kaixiang, et al.
Publicado: (2026)
por: Wang, Kaixiang, et al.
Publicado: (2026)
FedPFT: Federated Proxy Fine-Tuning of Foundation Models
por: Peng, Zhaopeng, et al.
Publicado: (2024)
por: Peng, Zhaopeng, et al.
Publicado: (2024)
CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM
por: Li, Yubo, et al.
Publicado: (2026)
por: Li, Yubo, et al.
Publicado: (2026)
Proxy-Based Approximation of Shapley and Banzhaf Interactions
por: Thies, Santo M. A. R., et al.
Publicado: (2026)
por: Thies, Santo M. A. R., et al.
Publicado: (2026)
CDKT-FL: Cross-Device Knowledge Transfer using Proxy Dataset in Federated Learning
por: Le, Huy Q., et al.
Publicado: (2022)
por: Le, Huy Q., et al.
Publicado: (2022)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
por: Chen, Ruishuo, et al.
Publicado: (2025)
por: Chen, Ruishuo, et al.
Publicado: (2025)
Transferring Causal Effects using Proxies
por: Iglesias-Alonso, Manuel, et al.
Publicado: (2025)
por: Iglesias-Alonso, Manuel, et al.
Publicado: (2025)
Constraint-Informed Active Learning for End-to-End ACOPF Optimization Proxies
por: Li, Miao, et al.
Publicado: (2025)
por: Li, Miao, et al.
Publicado: (2025)
Efficient Low Rank Attention for Long-Context Inference in Large Language Models
por: Li, Tenghui, et al.
Publicado: (2025)
por: Li, Tenghui, et al.
Publicado: (2025)
SPA-Cache: Singular Proxies for Adaptive Caching in Diffusion Language Models
por: Sun, Wenhao, et al.
Publicado: (2026)
por: Sun, Wenhao, et al.
Publicado: (2026)
TG-NAS: Generalizable Zero-Cost Proxies with Operator Description Embedding and Graph Learning for Efficient Neural Architecture Search
por: Qiao, Ye, et al.
Publicado: (2024)
por: Qiao, Ye, et al.
Publicado: (2024)
SentenceKV: Efficient LLM Inference via Sentence-Level Semantic KV Caching
por: Zhu, Yuxuan, et al.
Publicado: (2025)
por: Zhu, Yuxuan, et al.
Publicado: (2025)
KeepKV: Achieving Periodic Lossless KV Cache Compression for Efficient LLM Inference
por: Tian, Yuxuan, et al.
Publicado: (2025)
por: Tian, Yuxuan, et al.
Publicado: (2025)
ProxiMix: Enhancing Fairness with Proximity Samples in Subgroups
por: Hu, Jingyu, et al.
Publicado: (2024)
por: Hu, Jingyu, et al.
Publicado: (2024)
CSKV: Training-Efficient Channel Shrinking for KV Cache in Long-Context Scenarios
por: Wang, Luning, et al.
Publicado: (2024)
por: Wang, Luning, et al.
Publicado: (2024)
LaCache: Ladder-Shaped KV Caching for Efficient Long-Context Modeling of Large Language Models
por: Shi, Dachuan, et al.
Publicado: (2025)
por: Shi, Dachuan, et al.
Publicado: (2025)
Rethinking the Role of Proxy Rewards in Language Model Alignment
por: Kim, Sungdong, et al.
Publicado: (2024)
por: Kim, Sungdong, et al.
Publicado: (2024)
KV-Fold: One-Step KV-Cache Recurrence for Long-Context Inference
por: Nadali, Alireza, et al.
Publicado: (2026)
por: Nadali, Alireza, et al.
Publicado: (2026)
KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
por: Chen, Chuangtao, et al.
Publicado: (2026)
por: Chen, Chuangtao, et al.
Publicado: (2026)
Ejemplares similares
-
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
por: Li, Miao, et al.
Publicado: (2026) -
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
por: Zhu, Yu, et al.
Publicado: (2024) -
Predicting LLM Reasoning Performance with Small Proxy Model
por: Koh, Woosung, et al.
Publicado: (2025) -
LazyLLM: Dynamic Token Pruning for Efficient Long Context LLM Inference
por: Fu, Qichen, et al.
Publicado: (2024) -
Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
por: Liu, Zhihao, et al.
Publicado: (2025)