Saved in:
| Main Authors: | Dong, Sixun, Hu, Juhua, Li, Steven, Wen, Wei, Qian, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.04929 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
by: Wang, Chenyu, et al.
Published: (2024)
by: Wang, Chenyu, et al.
Published: (2024)
SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Pruning All-Rounder: Rethinking and Improving Inference Efficiency for Large Vision Language Models
by: Suo, Wei, et al.
Published: (2024)
by: Suo, Wei, et al.
Published: (2024)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
by: Qian, Qi, et al.
Published: (2024)
by: Qian, Qi, et al.
Published: (2024)
Text-Guided Mixup Towards Long-Tailed Image Categorization
by: Franklin, Richard, et al.
Published: (2024)
by: Franklin, Richard, et al.
Published: (2024)
Dual-disentangled Deep Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
Hierarchy-Aware Fine-Tuning of Vision-Language Models
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
by: Chen, Dong, et al.
Published: (2026)
by: Chen, Dong, et al.
Published: (2026)
Rethink Sparse Signals for Pose-guided Text-to-image Generation
by: Xuan, Wenjie, et al.
Published: (2025)
by: Xuan, Wenjie, et al.
Published: (2025)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
Rethinking Diffusion Model for Multi-Contrast MRI Super-Resolution
by: Li, Guangyuan, et al.
Published: (2024)
by: Li, Guangyuan, et al.
Published: (2024)
LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
by: Ye, Maoyuan, et al.
Published: (2025)
by: Ye, Maoyuan, et al.
Published: (2025)
Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference
by: Zhao, Han, et al.
Published: (2024)
by: Zhao, Han, et al.
Published: (2024)
Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
by: He, Haibin, et al.
Published: (2025)
by: He, Haibin, et al.
Published: (2025)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
by: Yuan, Qianhao, et al.
Published: (2025)
by: Yuan, Qianhao, et al.
Published: (2025)
Decoupled Multi-Predictor Optimization for Inference-Efficient Model Tuning
by: Luo, Liwei, et al.
Published: (2025)
by: Luo, Liwei, et al.
Published: (2025)
Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
Adapting Segment Anything Model for Power Transmission Corridor Hazard Segmentation
by: Chen, Hang, et al.
Published: (2025)
by: Chen, Hang, et al.
Published: (2025)
Rethinking Preference Alignment for Diffusion Models with Classifier-Free Guidance
by: Jiang, Zhou, et al.
Published: (2026)
by: Jiang, Zhou, et al.
Published: (2026)
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
by: Wu, Zongyu, et al.
Published: (2025)
by: Wu, Zongyu, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
by: Zhang, Ce, et al.
Published: (2025)
by: Zhang, Ce, et al.
Published: (2025)
Enhancing Consistency Models for Multi-Agent Trajectory Prediction
by: Mrdovic, Alen, et al.
Published: (2026)
by: Mrdovic, Alen, et al.
Published: (2026)
DiT-Air: Revisiting the Efficiency of Diffusion Model Architecture Design in Text to Image Generation
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Rethinking the Efficiency and Effectiveness of Reinforcement Learning for Radiology Report Generation
by: Lu, Zilin, et al.
Published: (2026)
by: Lu, Zilin, et al.
Published: (2026)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
Multi-modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation
by: Li, Xiawei, et al.
Published: (2023)
by: Li, Xiawei, et al.
Published: (2023)
Rethinking the Need for Source Models: Source-Free Domain Adaptation from Scratch Guided by a Vision-Language Model
by: Bingtao, Zhou, et al.
Published: (2026)
by: Bingtao, Zhou, et al.
Published: (2026)
LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models
by: Liu, Lu, et al.
Published: (2026)
by: Liu, Lu, et al.
Published: (2026)
Rethinking Model Ensemble in Transfer-based Adversarial Attacks
by: Chen, Huanran, et al.
Published: (2023)
by: Chen, Huanran, et al.
Published: (2023)
ArtiCAD: Articulated CAD Assembly Design via Multi-Agent Code Generation
by: Shui, Yuan, et al.
Published: (2026)
by: Shui, Yuan, et al.
Published: (2026)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
by: Qian, Kun, et al.
Published: (2025)
by: Qian, Kun, et al.
Published: (2025)
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
Understanding Multi-Agent Reasoning with Large Language Models for Cartoon VQA
by: Wu, Tong, et al.
Published: (2026)
by: Wu, Tong, et al.
Published: (2026)
Similar Items
-
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
by: Dong, Sixun, et al.
Published: (2025) -
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024) -
Online Zero-Shot Classification with CLIP
by: Qian, Qi, et al.
Published: (2024) -
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
by: Wang, Chenyu, et al.
Published: (2024) -
SeA: Semantic Adversarial Augmentation for Last Layer Features from Unsupervised Representation Learning
by: Qian, Qi, et al.
Published: (2024)