ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ouyang, Haojie, Lv, Jianwei, Ren, Lei, Wei, Chen, Wang, Xiaojie, Feng, Fangxiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities
von: Xia, Guoyang, et al.
Veröffentlicht: (2025)
von: Xia, Guoyang, et al.
Veröffentlicht: (2025)
A Pluggable Multi-Task Learning Framework for Sentiment-Aware Financial Relation Extraction
von: Luo, Jinming, et al.
Veröffentlicht: (2025)
von: Luo, Jinming, et al.
Veröffentlicht: (2025)
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
von: Tang, Yaohua, et al.
Veröffentlicht: (2025)
von: Tang, Yaohua, et al.
Veröffentlicht: (2025)
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
von: Wang, Jiahao, et al.
Veröffentlicht: (2026)
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
QCG-Rerank: Chunks Graph Rerank with Query Expansion in Retrieval-Augmented LLMs for Tourism Domain
von: Wei, Qikai, et al.
Veröffentlicht: (2024)
von: Wei, Qikai, et al.
Veröffentlicht: (2024)
HiChunk: Evaluating and Enhancing Retrieval-Augmented Generation with Hierarchical Chunking
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
von: Lu, Wensheng, et al.
Veröffentlicht: (2025)
Provable Defense Framework for LLM Jailbreaks via Noise-Augumented Alignment
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
von: Cheng, Zehua, et al.
Veröffentlicht: (2026)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
von: Li, Haoyang, et al.
Veröffentlicht: (2025)
A Lightweight Framework for Trigger-Guided LoRA-Based Self-Adaptation in LLMs
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
von: Wei, Jiacheng, et al.
Veröffentlicht: (2025)
Adaptive Token Boundaries: Integrating Human Chunking Mechanisms into Multimodal LLMs
von: Yu, Dongxing
Veröffentlicht: (2025)
von: Yu, Dongxing
Veröffentlicht: (2025)
A Pluggable Common Sense-Enhanced Framework for Knowledge Graph Completion
von: Niu, Guanglin, et al.
Veröffentlicht: (2024)
von: Niu, Guanglin, et al.
Veröffentlicht: (2024)
SampleAttention: Near-Lossless Acceleration of Long Context LLM Inference with Adaptive Structured Sparse Attention
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
von: Zhu, Qianchao, et al.
Veröffentlicht: (2024)
H-PRM: A Pluggable Hotword Pre-Retrieval Module for Various Speech Recognition Systems
von: Dai, Huangyu, et al.
Veröffentlicht: (2025)
von: Dai, Huangyu, et al.
Veröffentlicht: (2025)
The Impact of Inference Acceleration on Bias of LLMs
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
von: Kirsten, Elisabeth, et al.
Veröffentlicht: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
von: Wen, Zhuofan, et al.
Veröffentlicht: (2024)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
von: Feng, Yuan, et al.
Veröffentlicht: (2024)
Clover-2: Accurate Inference for Regressive Lightweight Speculative Decoding
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
von: Xiao, Bin, et al.
Veröffentlicht: (2024)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
von: Duan, Shaohua, et al.
Veröffentlicht: (2025)
von: Duan, Shaohua, et al.
Veröffentlicht: (2025)
LycheeDecode: Accelerating Long-Context LLM Inference via Hybrid-Head Sparse Decoding
von: Lin, Gang, et al.
Veröffentlicht: (2026)
von: Lin, Gang, et al.
Veröffentlicht: (2026)
Concept Conductor: Orchestrating Multiple Personalized Concepts in Text-to-Image Synthesis
von: Yao, Zebin, et al.
Veröffentlicht: (2024)
von: Yao, Zebin, et al.
Veröffentlicht: (2024)
Chunking German Legal Code
von: Prior, Max, et al.
Veröffentlicht: (2026)
von: Prior, Max, et al.
Veröffentlicht: (2026)
Chunk-Distilled Language Modeling
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
von: Li, Yanhong, et al.
Veröffentlicht: (2024)
Amphista: Bi-directional Multi-head Decoding for Accelerating LLM Inference
von: Li, Zeping, et al.
Veröffentlicht: (2024)
von: Li, Zeping, et al.
Veröffentlicht: (2024)
Scaling LLM Inference with Optimized Sample Compute Allocation
von: Zhang, Kexun, et al.
Veröffentlicht: (2024)
von: Zhang, Kexun, et al.
Veröffentlicht: (2024)
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
von: Júnior, Paulo Roberto de Moura, et al.
Veröffentlicht: (2026)
von: Júnior, Paulo Roberto de Moura, et al.
Veröffentlicht: (2026)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
von: He, Zhuomin, et al.
Veröffentlicht: (2025)
LLM-based Translation Inference with Iterative Bilingual Understanding
von: Chen, Andong, et al.
Veröffentlicht: (2024)
von: Chen, Andong, et al.
Veröffentlicht: (2024)
FASST: Fast LLM-based Simultaneous Speech Translation
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
LLM Inference Acceleration via Efficient Operation Fusion
von: Salmani, Mahsa, et al.
Veröffentlicht: (2025)
von: Salmani, Mahsa, et al.
Veröffentlicht: (2025)
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
von: Xiao, Qingfa, et al.
Veröffentlicht: (2025)
CritiPrefill: A Segment-wise Criticality-based Approach for Prefilling Acceleration in LLMs
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
von: Lv, Junlin, et al.
Veröffentlicht: (2024)
Language Ranker: A Lightweight Ranking framework for LLM Decoding
von: Zhang, Chenheng, et al.
Veröffentlicht: (2025)
von: Zhang, Chenheng, et al.
Veröffentlicht: (2025)
LLMs Are Prone to Fallacies in Causal Inference
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
von: Joshi, Nitish, et al.
Veröffentlicht: (2024)
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Feng, Zhaopeng, et al.
Veröffentlicht: (2025)
Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
von: Li, Jianwei, et al.
Veröffentlicht: (2023)
A Lightweight LLM Framework for Disaster Humanitarian Information Classification
von: Jinzhen, Han, et al.
Veröffentlicht: (2026)
von: Jinzhen, Han, et al.
Veröffentlicht: (2026)
GradPruner: Gradient-Guided Layer Pruning Enabling Efficient Fine-Tuning and Inference for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic Language
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
XL3M: A Training-free Framework for LLM Length Extension Based on Segment-wise Inference
von: Wang, Shengnan, et al.
Veröffentlicht: (2024)
von: Wang, Shengnan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities
von: Xia, Guoyang, et al.
Veröffentlicht: (2025) -
A Pluggable Multi-Task Learning Framework for Sentiment-Aware Financial Relation Extraction
von: Luo, Jinming, et al.
Veröffentlicht: (2025) -
Round Attention: A Novel Round-Level Attention Mechanism to Accelerate LLM Inference
von: Tang, Yaohua, et al.
Veröffentlicht: (2025) -
From Prefix Cache to Fusion RAG Cache: Accelerating LLM Inference in Retrieval-Augmented Generation
von: Wang, Jiahao, et al.
Veröffentlicht: (2026) -
Cascaded Self-Evaluation Augmented Training for Lightweight Multimodal LLMs
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)