QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhongyang, Li, Yaqian, Fang, Faming, Takezoe, Rinyoichi, Bo, Zi-Hao, Qian, Cheng, Guang, Mo, Zhang, Guixu, Long, Kaiwen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
by: Takezoe, Rinyoichi, et al.
Published: (2026)
by: Takezoe, Rinyoichi, et al.
Published: (2026)
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
by: Bo, Zi-Hao, et al.
Published: (2026)
by: Bo, Zi-Hao, et al.
Published: (2026)
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
by: Zhao, Yanxiao, et al.
Published: (2025)
by: Zhao, Yanxiao, et al.
Published: (2025)
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
by: Liu, Hanpeng, et al.
Published: (2026)
by: Liu, Hanpeng, et al.
Published: (2026)
Deep Unfolding Convolutional Dictionary Model for Multi-Contrast MRI Super-resolution and Reconstruction
by: Lei, Pengcheng, et al.
Published: (2023)
by: Lei, Pengcheng, et al.
Published: (2023)
First-order State Space Model for Lightweight Image Super-resolution
by: Zhu, Yujie, et al.
Published: (2025)
by: Zhu, Yujie, et al.
Published: (2025)
TokenPacker: Efficient Visual Projector for Multimodal LLM
by: Li, Wentong, et al.
Published: (2024)
by: Li, Wentong, et al.
Published: (2024)
Harmonizing knowledge Transfer in Neural Network with Unified Distillation
by: Huang, Yaomin, et al.
Published: (2024)
by: Huang, Yaomin, et al.
Published: (2024)
Tiny-QMoE
by: Cashman, Jack, et al.
Published: (2025)
by: Cashman, Jack, et al.
Published: (2025)
QMoE: A Quantum Mixture of Experts Framework for Scalable Quantum Neural Networks
by: Nguyen, Hoang-Quan, et al.
Published: (2025)
by: Nguyen, Hoang-Quan, et al.
Published: (2025)
Exact Recovery of Community Detection in dependent Gaussian Mixture Models
by: Li, Zhongyang, et al.
Published: (2022)
by: Li, Zhongyang, et al.
Published: (2022)
CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
by: Dong, Jiale, et al.
Published: (2025)
by: Dong, Jiale, et al.
Published: (2025)
Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation
by: Gao, Juntao, et al.
Published: (2025)
by: Gao, Juntao, et al.
Published: (2025)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
by: Li, Shuai, et al.
Published: (2025)
by: Li, Shuai, et al.
Published: (2025)
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
R2-T2: Re-Routing in Test-Time for Multimodal Mixture-of-Experts
by: Li, Zhongyang, et al.
Published: (2025)
by: Li, Zhongyang, et al.
Published: (2025)
Frequency Error-Guided Under-sampling Optimization for Multi-Contrast MRI Reconstruction
by: Fang, Xinming, et al.
Published: (2026)
by: Fang, Xinming, et al.
Published: (2026)
An Efficient Token Compression Framework for Visual Object Tracking
by: Wu, Weijing, et al.
Published: (2026)
by: Wu, Weijing, et al.
Published: (2026)
Monocular Depth Estimation with Global-Aware Discretization and Local Context Modeling
by: Wu, Heng, et al.
Published: (2025)
by: Wu, Heng, et al.
Published: (2025)
LRCP: Low-Rank Compressibility Guided Visual Token Pruning for Efficient LVLMs
by: Lu, Hongyu, et al.
Published: (2026)
by: Lu, Hongyu, et al.
Published: (2026)
Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
by: Peng, Tianfan, et al.
Published: (2025)
by: Peng, Tianfan, et al.
Published: (2025)
Stationary properties of the Gauss-Galerkin QMoM truncation of MV-SDEs
by: Alecio, Alexander
Published: (2024)
by: Alecio, Alexander
Published: (2024)
RUL-QMoE: Multiple Non-crossing Quantile Mixture-of-Experts for Probabilistic Remaining Useful Life Predictions of Varying Battery Materials
by: Ly, Sel, et al.
Published: (2025)
by: Ly, Sel, et al.
Published: (2025)
OneLatent: Single-Token Compression for Visual Latent Reasoning
by: Lv, Bo, et al.
Published: (2026)
by: Lv, Bo, et al.
Published: (2026)
The Security Threat of Compressed Projectors in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2025)
by: Zhang, Yudong, et al.
Published: (2025)
A 160 ° x 160 ° Dynamic Holographic Meta-Projector
by: Li, Feng-Jun, et al.
Published: (2025)
by: Li, Feng-Jun, et al.
Published: (2025)
Definition of CQL, a Visual Query Language
by: Shi-Guang Ju
Published: (1999)
by: Shi-Guang Ju
Published: (1999)
Matchings on Random Regular Hypergraphs
by: Li, Zhongyang
Published: (2021)
by: Li, Zhongyang
Published: (2021)
Planar Site Percolation, End Structure, and the Benjamini-Schramm Conjecture
by: Li, Zhongyang
Published: (2026)
by: Li, Zhongyang
Published: (2026)
Perfect Matchings and Essential Spanning Forests in Hyperbolic Double Circle Packings
by: Li, Zhongyang
Published: (2024)
by: Li, Zhongyang
Published: (2024)
Recursive Packing Bounds for Supercritical Disconnection in Bernoulli Site Percolation
by: Li, Zhongyang
Published: (2026)
by: Li, Zhongyang
Published: (2026)
Independent GUE minor processes of perfect matchings on rail-yard graphs
by: Li, Zhongyang
Published: (2024)
by: Li, Zhongyang
Published: (2024)
Critical site percolation and cutsets
by: Li, Zhongyang
Published: (2024)
by: Li, Zhongyang
Published: (2024)
Tree embeddings and nonuniqueness in site percolation
by: Li, Zhongyang
Published: (2023)
by: Li, Zhongyang
Published: (2023)
C‐TUnet: A CNN‐Transformer Architecture‐Based Ultrasound Breast Image Classification Network
by: Ying Wu, et al.
Published: (2024)
by: Ying Wu, et al.
Published: (2024)
VOLoc: Visual Place Recognition by Querying Compressed Lidar Map
by: Cai, Xudong, et al.
Published: (2024)
by: Cai, Xudong, et al.
Published: (2024)
Setup-Independent Full Projector Compensation
by: Li, Haibo, et al.
Published: (2026)
by: Li, Haibo, et al.
Published: (2026)
AudioMoG: Guiding Audio Generation with Mixture-of-Guidance
by: Wang, Junyou, et al.
Published: (2025)
by: Wang, Junyou, et al.
Published: (2025)
Visual-Guided Key-Token Regularization for Multimodal Large Language Model Unlearning
by: Cai, Chengyi, et al.
Published: (2026)
by: Cai, Chengyi, et al.
Published: (2026)
Similar Items
-
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
by: Takezoe, Rinyoichi, et al.
Published: (2026) -
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
by: Bo, Zi-Hao, et al.
Published: (2026) -
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
by: Zhao, Yanxiao, et al.
Published: (2025) -
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
by: Liu, Hanpeng, et al.
Published: (2026) -
ITO: Images and Texts as One via Synergizing Multiple Alignment and Training-Time Fusion
by: Liu, Hanpeng, et al.
Published: (2026)