CAMERA: Multi-Matrix Joint Compression for MoE Models via Micro-Expert Redundancy Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Yuzhuang, Han, Xu, Zhang, Yuanchi, Wang, Yixuan, Liu, Yijun, Ji, Shiyu, Zhu, Qingfu, Che, Wanxiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025)
by: Liu, Yijun, et al.
Published: (2025)
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
by: Ji, Shiyu, et al.
Published: (2026)
by: Ji, Shiyu, et al.
Published: (2026)
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Seer Self-Consistency: Advance Budget Estimation for Adaptive Test-Time Scaling
by: Ji, Shiyu, et al.
Published: (2025)
by: Ji, Shiyu, et al.
Published: (2025)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization
by: Shan, Baocai, et al.
Published: (2026)
by: Shan, Baocai, et al.
Published: (2026)
Improving Grammatical Error Correction via Contextual Data Augmentation
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
CommonKV: Compressing KV Cache with Cross-layer Parameter Sharing
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
OneBit: Towards Extremely Low-bit Large Language Models
by: Xu, Yuzhuang, et al.
Published: (2024)
by: Xu, Yuzhuang, et al.
Published: (2024)
Fitting Is Not Enough: Smoothness in Extremely Quantized LLMs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
by: Zhang, Geng, et al.
Published: (2025)
by: Zhang, Geng, et al.
Published: (2025)
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring
by: Niu, Tianhao, et al.
Published: (2026)
by: Niu, Tianhao, et al.
Published: (2026)
Sub-MoE: Efficient Mixture-of-Expert LLMs Compression via Subspace Expert Merging
by: Li, Lujun, et al.
Published: (2025)
by: Li, Lujun, et al.
Published: (2025)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
by: Ma, Songkai, et al.
Published: (2025)
by: Ma, Songkai, et al.
Published: (2025)
How Do Language Models Understand Tables? A Mechanistic Analysis of Cell Location
by: Zhang, Xuanliang, et al.
Published: (2026)
by: Zhang, Xuanliang, et al.
Published: (2026)
Scaling Laws for Agent Harnesses via Effective Feedback Compute
by: Zhang, Xuanliang, et al.
Published: (2026)
by: Zhang, Xuanliang, et al.
Published: (2026)
MoE-Prefill: Zero Redundancy Overheads in MoE Prefill Serving
by: Su, Zhaoyuan, et al.
Published: (2026)
by: Su, Zhaoyuan, et al.
Published: (2026)
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
by: Neogi, Pinaki Prasad Guha, et al.
Published: (2025)
by: Neogi, Pinaki Prasad Guha, et al.
Published: (2025)
MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs
by: Chen, Xiaodong, et al.
Published: (2025)
by: Chen, Xiaodong, et al.
Published: (2025)
MergeMoE: Efficient Compression of MoE Models via Expert Output Merging
by: Miao, Ruijie, et al.
Published: (2025)
by: Miao, Ruijie, et al.
Published: (2025)
Make Some Noise: Unlocking Language Model Parallel Inference Capability through Noisy Training
by: Wang, Yixuan, et al.
Published: (2024)
by: Wang, Yixuan, et al.
Published: (2024)
Multi-Layer Attention is the Amplifier of Demonstration Effectiveness
by: Wang, Dingzirui, et al.
Published: (2025)
by: Wang, Dingzirui, et al.
Published: (2025)
R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning
by: Guo, Xiaohan, et al.
Published: (2025)
by: Guo, Xiaohan, et al.
Published: (2025)
EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference
by: Qian, Yulei, et al.
Published: (2024)
by: Qian, Yulei, et al.
Published: (2024)
Abacus-SQL: A Text-to-SQL System Empowering Cross-Domain and Open-Domain Database Retrieval
by: Xu, Keyan, et al.
Published: (2025)
by: Xu, Keyan, et al.
Published: (2025)
RoT: Enhancing Table Reasoning with Iterative Row-Wise Traversals
by: Zhang, Xuanliang, et al.
Published: (2025)
by: Zhang, Xuanliang, et al.
Published: (2025)
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering
by: Zhang, Xuanliang, et al.
Published: (2025)
by: Zhang, Xuanliang, et al.
Published: (2025)
Advancing Expert Specialization for Better MoE
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
by: Li, Zichong, et al.
Published: (2025)
by: Li, Zichong, et al.
Published: (2025)
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
by: Zhu, Peijun, et al.
Published: (2025)
by: Zhu, Peijun, et al.
Published: (2025)
How Many Code and Test Cases Are Enough? Evaluating Test Cases Generation from a Binary-Matrix Perspective
by: Luo, Xianzhen, et al.
Published: (2025)
by: Luo, Xianzhen, et al.
Published: (2025)
ProxyAttn: Guided Sparse Attention via Representative Heads
by: Wang, Yixuan, et al.
Published: (2025)
by: Wang, Yixuan, et al.
Published: (2025)
Leave It to the Experts: Detecting Knowledge Distillation via MoE Expert Signatures
by: Li, Pingzhi, et al.
Published: (2025)
by: Li, Pingzhi, et al.
Published: (2025)
Automated Snippet-Alignment Data Augmentation for Code Translation
by: Zhang, Zhiming, et al.
Published: (2025)
by: Zhang, Zhiming, et al.
Published: (2025)
MURRE: Multi-Hop Table Retrieval with Removal for Open-Domain Text-to-SQL
by: Zhang, Xuanliang, et al.
Published: (2024)
by: Zhang, Xuanliang, et al.
Published: (2024)
PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts into Prompt Tuning
by: Li, Zongqian, et al.
Published: (2025)
by: Li, Zongqian, et al.
Published: (2025)
$μ$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
by: Koike-Akino, Toshiaki, et al.
Published: (2025)
Exploring Hybrid Question Answering via Program-based Prompting
by: Shi, Qi, et al.
Published: (2024)
by: Shi, Qi, et al.
Published: (2024)
Similar Items
-
CRVQ: Channel-Relaxed Vector Quantization for Extreme Compression of LLMs
by: Xu, Yuzhuang, et al.
Published: (2024) -
Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
by: Wang, Yixuan, et al.
Published: (2025) -
Judge Q: Trainable Queries for Optimized Information Retention in KV Cache Eviction
by: Liu, Yijun, et al.
Published: (2025) -
EchoKV: Efficient KV Cache Compression via Similarity-Based Reconstruction
by: Ji, Shiyu, et al.
Published: (2026) -
Think Before You Accept: Semantic Reflective Verification for Faster Speculative Decoding
by: Wang, Yixuan, et al.
Published: (2025)