Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zeliang, Pham, Phu, Zhao, Wentian, Wan, Kun, Li, Yu-Jhe, Zhou, Jianing, Miranda, Daniel, Kale, Ajinkya, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
by: Deng, Shijian, et al.
Published: (2024)
by: Deng, Shijian, et al.
Published: (2024)
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
by: Li, Yu-Jhe, et al.
Published: (2024)
by: Li, Yu-Jhe, et al.
Published: (2024)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025)
by: Wang, Zhenting, et al.
Published: (2025)
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024)
by: Li, Kevin Y., et al.
Published: (2024)
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
See the Text: From Tokenization to Visual Reading
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
Approximated Likelihood Ratio: A Forward-Only and Parallel Framework for Boosting Neural Network Training
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
DRIFT: Transferring Reasoning Priors for Efficient MLLM Fine-Tuning
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs
by: Han, Insu, et al.
Published: (2025)
by: Han, Insu, et al.
Published: (2025)
An Attention‐Driven Graph Transformer With Nonlinear Modeling and Neuro‐Fuzzy Fusion for High‐Order Toxic Molecular Graph Learning
by: Phu Pham
Published: (2026)
by: Phu Pham
Published: (2026)
Recognize How Your Marketing Efforts May Need a Change
by: Alison Knopf
Published: (2024)
by: Alison Knopf
Published: (2024)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
Discover and Mitigate Multiple Biased Subgroups in Image Classifiers
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
by: Chen, Shimin, et al.
Published: (2024)
by: Chen, Shimin, et al.
Published: (2024)
Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
by: Chi, Donghwan, et al.
Published: (2025)
by: Chi, Donghwan, et al.
Published: (2025)
Early-Stage Prediction of Review Effort in AI-Generated Pull Requests
by: Minh, Dao Sy Duy, et al.
Published: (2026)
by: Minh, Dao Sy Duy, et al.
Published: (2026)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
Fewer Tokens, Greater Scaling: Self-Adaptive Visual Bases for Efficient and Expansive Representation Learning
by: Young, Shawn, et al.
Published: (2025)
by: Young, Shawn, et al.
Published: (2025)
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
by: Zhao, Guangxiang, et al.
Published: (2026)
by: Zhao, Guangxiang, et al.
Published: (2026)
EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning
by: Xu, Wujiang, et al.
Published: (2025)
by: Xu, Wujiang, et al.
Published: (2025)
ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
by: Huang, Runhui, et al.
Published: (2025)
by: Huang, Runhui, et al.
Published: (2025)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
The Only Copyright Law We Need.
by: Toohey, Daniel
Published: (1984)
by: Toohey, Daniel
Published: (1984)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
by: Brack, Manuel, et al.
Published: (2025)
by: Brack, Manuel, et al.
Published: (2025)
Optimizing Crowd-Aware Multi-Agent Path Finding through Local Communication with Graph Neural Networks
by: Pham, Phu, et al.
Published: (2023)
by: Pham, Phu, et al.
Published: (2023)
Integrating Corpus Analysis and ChatGPT in Teaching English Collocations: A Hybrid Approach
by: Quy Huynh Phu Pham
Published: (2025)
by: Quy Huynh Phu Pham
Published: (2025)
Understanding and Mitigating Numerical Sources of Nondeterminism in LLM Inference
by: Yuan, Jiayi, et al.
Published: (2025)
by: Yuan, Jiayi, et al.
Published: (2025)
Token Coordinated Prompt Attention is Needed for Visual Prompting
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
by: Hu, Lianyu, et al.
Published: (2024)
by: Hu, Lianyu, et al.
Published: (2024)
Autoregressive Image Generation Needs Only a Few Lines of Cached Tokens
by: Qin, Ziran, et al.
Published: (2025)
by: Qin, Ziran, et al.
Published: (2025)
Learning to Transform Dynamically for Better Adversarial Transferability
by: Zhu, Rongyi, et al.
Published: (2024)
by: Zhu, Rongyi, et al.
Published: (2024)
Forward Learning with Differential Privacy
by: Feng, Mingqian, et al.
Published: (2025)
by: Feng, Mingqian, et al.
Published: (2025)
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Why Instruction-Based Unlearning Fails in Diffusion Models?
by: Zhang, Zeliang, et al.
Published: (2026)
by: Zhang, Zeliang, et al.
Published: (2026)
Will the Inclusion of Generated Data Amplify Bias Across Generations in Future Image Classification Models?
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
Similar Items
-
Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach
by: Deng, Shijian, et al.
Published: (2024) -
Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation
by: Li, Yu-Jhe, et al.
Published: (2024) -
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
by: Wang, Zhenting, et al.
Published: (2025) -
Inference Optimal VLMs Need Fewer Visual Tokens and More Parameters
by: Li, Kevin Y., et al.
Published: (2024) -
OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answering
by: Vosoughi, Ali, et al.
Published: (2025)