Saved in:
| Main Authors: | Tao, Wei, Zhang, Bin, Qu, Xiaoyang, Wan, Jiguang, Wang, Jianzong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2503.23294 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026)
by: Tao, Wei, et al.
Published: (2026)
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
by: Tao, Wei, et al.
Published: (2024)
by: Tao, Wei, et al.
Published: (2024)
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic Segmentation
by: Tao, Wei, et al.
Published: (2025)
by: Tao, Wei, et al.
Published: (2025)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
by: Wang, Anmin, et al.
Published: (2026)
by: Wang, Anmin, et al.
Published: (2026)
Vista: Scene-Aware Optimization for Streaming Video Question Answering under Post-Hoc Queries
by: Lu, Haocheng, et al.
Published: (2026)
by: Lu, Haocheng, et al.
Published: (2026)
PQCache: Product Quantization-based KVCache for Long Context LLM Inference
by: Zhang, Hailin, et al.
Published: (2024)
by: Zhang, Hailin, et al.
Published: (2024)
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
MiTa: A Hierarchical Multi-Agent Collaboration Framework with Memory-integrated and Task Allocation
by: Zhang, XiaoJie, et al.
Published: (2026)
by: Zhang, XiaoJie, et al.
Published: (2026)
PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action Recognition
by: He, Shenglin, et al.
Published: (2024)
by: He, Shenglin, et al.
Published: (2024)
AdaSkip: Adaptive Sublayer Skipping for Accelerating Long-Context LLM Inference
by: He, Zhuomin, et al.
Published: (2025)
by: He, Zhuomin, et al.
Published: (2025)
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
by: Li, Xing, et al.
Published: (2025)
by: Li, Xing, et al.
Published: (2025)
RATE-Nav: Region-Aware Termination Enhancement for Zero-shot Object Navigation with Vision-Language Models
by: Li, Junjie, et al.
Published: (2025)
by: Li, Junjie, et al.
Published: (2025)
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
by: Ouyang, Haojie, et al.
Published: (2025)
by: Ouyang, Haojie, et al.
Published: (2025)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
by: Li, Dongfang, et al.
Published: (2026)
by: Li, Dongfang, et al.
Published: (2026)
The Cambrian Explosion of Mixed-Precision Matrix Multiplication for Quantized Deep Learning Inference
by: Martínez, Héctor, et al.
Published: (2025)
by: Martínez, Héctor, et al.
Published: (2025)
Late Chunking: Contextual Chunk Embeddings Using Long-Context Embedding Models
by: Günther, Michael, et al.
Published: (2024)
by: Günther, Michael, et al.
Published: (2024)
Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization
by: Duan, Shaohua, et al.
Published: (2025)
by: Duan, Shaohua, et al.
Published: (2025)
SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
FocusLLM: Precise Understanding of Long Context by Dynamic Condensing
by: Li, Zhenyu, et al.
Published: (2024)
by: Li, Zhenyu, et al.
Published: (2024)
Progressive Mixed-Precision Decoding for Efficient LLM Inference
by: Chen, Hao Mark, et al.
Published: (2024)
by: Chen, Hao Mark, et al.
Published: (2024)
Dynamic Chunking and Selection for Reading Comprehension of Ultra-Long Context in Large Language Models
by: Sheng, Boheng, et al.
Published: (2025)
by: Sheng, Boheng, et al.
Published: (2025)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
by: Zhou, Lang, et al.
Published: (2026)
by: Zhou, Lang, et al.
Published: (2026)
Efficient Long-Context LLM Inference via KV Cache Clustering
by: Hu, Jie, et al.
Published: (2025)
by: Hu, Jie, et al.
Published: (2025)
ChunkWise LoRA: Adaptive Sequence Partitioning for Memory-Efficient Low-Rank Adaptation and Accelerated LLM Inference
by: Thakkar, Ketan, et al.
Published: (2026)
by: Thakkar, Ketan, et al.
Published: (2026)
Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
by: Lu, Haiquan, et al.
Published: (2026)
by: Lu, Haiquan, et al.
Published: (2026)
MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference
by: Wan, Zhongwei, et al.
Published: (2025)
by: Wan, Zhongwei, et al.
Published: (2025)
QAQ: Quality Adaptive Quantization for LLM KV Cache
by: Dong, Shichen, et al.
Published: (2024)
by: Dong, Shichen, et al.
Published: (2024)
D2O: Dynamic Discriminative Operations for Efficient Long-Context Inference of Large Language Models
by: Wan, Zhongwei, et al.
Published: (2024)
by: Wan, Zhongwei, et al.
Published: (2024)
Context-Aware Dynamic Chunking for Streaming Tibetan Speech Recognition
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context Focusing
by: Long, Lingkun, et al.
Published: (2026)
by: Long, Lingkun, et al.
Published: (2026)
Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-based Planner and Graph-based Policy
by: Jia, Ziqi, et al.
Published: (2025)
by: Jia, Ziqi, et al.
Published: (2025)
Adaptive Chunking: Optimizing Chunking-Method Selection for RAG
by: Júnior, Paulo Roberto de Moura, et al.
Published: (2026)
by: Júnior, Paulo Roberto de Moura, et al.
Published: (2026)
Squeezed Attention: Accelerating Long Context Length LLM Inference
by: Hooper, Coleman, et al.
Published: (2024)
by: Hooper, Coleman, et al.
Published: (2024)
TailorKV: A Hybrid Framework for Long-Context Inference via Tailored KV Cache Optimization
by: Yao, Dingyu, et al.
Published: (2025)
by: Yao, Dingyu, et al.
Published: (2025)
GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution Detection
by: Chen, Jinggang, et al.
Published: (2023)
by: Chen, Jinggang, et al.
Published: (2023)
Similar Items
-
MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware Experts
by: Tao, Wei, et al.
Published: (2025) -
WindowQuant: Mixed-Precision KV Cache Quantization based on Window-Level Similarity for VLMs Inference Optimization
by: Tao, Wei, et al.
Published: (2026) -
Value-Driven Mixed-Precision Quantization for Patch-Based Inference on Microcontrollers
by: Tao, Wei, et al.
Published: (2024) -
MADLLM: Multivariate Anomaly Detection via Pre-trained LLMs
by: Tao, Wei, et al.
Published: (2025) -
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
by: Zhang, Bin, et al.
Published: (2025)