Instruction Tuning-free Visual Token Complement for Multimodal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Dongsheng, Cui, Jiequan, Li, Miaoge, Lin, Wang, Chen, Bo, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness
von: Zhu, Beier, et al.
Veröffentlicht: (2025)
von: Zhu, Beier, et al.
Veröffentlicht: (2025)
Robust Fine-tuning of Zero-shot Models via Variance Reduction
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
von: Zhu, Beier, et al.
Veröffentlicht: (2024)
Dynamic Multimodal Prototype Learning in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
LoRA of Change: Learning to Generate LoRA for the Editing Instruction from A Single Before-After Image Pair
von: Song, Xue, et al.
Veröffentlicht: (2024)
von: Song, Xue, et al.
Veröffentlicht: (2024)
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)
Doubly Abductive Counterfactual Inference for Text-based Image Editing
von: Song, Xue, et al.
Veröffentlicht: (2024)
von: Song, Xue, et al.
Veröffentlicht: (2024)
Patch-Prompt Aligned Bayesian Prompt Tuning for Vision-Language Models
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
von: Liu, Xinyang, et al.
Veröffentlicht: (2023)
Pushing Rendering Boundaries: Hard Gaussian Splatting
von: Xu, Qingshan, et al.
Veröffentlicht: (2024)
von: Xu, Qingshan, et al.
Veröffentlicht: (2024)
STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
von: Li, Miaoge, et al.
Veröffentlicht: (2026)
von: Li, Miaoge, et al.
Veröffentlicht: (2026)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
Aligned Contrastive Loss for Long-Tailed Recognition
von: Ma, Jiali, et al.
Veröffentlicht: (2025)
von: Ma, Jiali, et al.
Veröffentlicht: (2025)
Classes Are Not Equal: An Empirical Study on Image Recognition Fairness
von: Cui, Jiequan, et al.
Veröffentlicht: (2024)
von: Cui, Jiequan, et al.
Veröffentlicht: (2024)
Decoupled Kullback-Leibler Divergence Loss
von: Cui, Jiequan, et al.
Veröffentlicht: (2023)
von: Cui, Jiequan, et al.
Veröffentlicht: (2023)
Generative Multimodal Pretraining with Discrete Diffusion Timestep Tokens
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Fine-tuning Multimodal LLMs to Follow Zero-shot Demonstrative Instructions
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
von: Li, Juncheng, et al.
Veröffentlicht: (2023)
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
von: li, Bonan, et al.
Veröffentlicht: (2025)
von: li, Bonan, et al.
Veröffentlicht: (2025)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024)
von: Li, Honglin, et al.
Veröffentlicht: (2024)
Visual Instruction Tuning with Chain of Region-of-Interest
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
Reducing Class-Wise Performance Disparity via Margin Regularization
von: Zhu, Beier, et al.
Veröffentlicht: (2026)
von: Zhu, Beier, et al.
Veröffentlicht: (2026)
NeuSpring: Neural Spring Fields for Reconstruction and Simulation of Deformable Objects from Videos
von: Xu, Qingshan, et al.
Veröffentlicht: (2025)
von: Xu, Qingshan, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
TrimTokenator-LC: Towards Adaptive Visual Token Pruning for Large Multimodal Models with Long Contexts
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Multimodal Instruction Tuning with Hybrid State Space Models
von: Zhou, Jianing, et al.
Veröffentlicht: (2024)
von: Zhou, Jianing, et al.
Veröffentlicht: (2024)
Token Activation Map to Visually Explain Multimodal LLMs
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
von: Li, Songze, et al.
Veröffentlicht: (2025)
von: Li, Songze, et al.
Veröffentlicht: (2025)
Feature Complementation Architecture for Visual Place Recognition
von: Wang, Weiwei, et al.
Veröffentlicht: (2025)
von: Wang, Weiwei, et al.
Veröffentlicht: (2025)
Auto-Encoding Morph-Tokens for Multimodal LLM
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
von: Wang, Bohan, et al.
Veröffentlicht: (2025)
von: Wang, Bohan, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
TsCA: On the Semantic Consistency Alignment via Conditional Transport for Compositional Zero-Shot Learning
von: Li, Miaoge, et al.
Veröffentlicht: (2024)
von: Li, Miaoge, et al.
Veröffentlicht: (2024)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
von: Guo, Jarvis, et al.
Veröffentlicht: (2024)
Exploring Transferable Homogeneous Groups for Compositional Zero-Shot Learning
von: Rao, Zhijie, et al.
Veröffentlicht: (2025)
von: Rao, Zhijie, et al.
Veröffentlicht: (2025)
End-to-End Vision Tokenizer Tuning
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
TokenPacker: Efficient Visual Projector for Multimodal LLM
von: Li, Wentong, et al.
Veröffentlicht: (2024)
von: Li, Wentong, et al.
Veröffentlicht: (2024)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
von: Du, Yifan, et al.
Veröffentlicht: (2023)
von: Du, Yifan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Project-Probe-Aggregate: Efficient Fine-Tuning for Group Robustness
von: Zhu, Beier, et al.
Veröffentlicht: (2025) -
Robust Fine-tuning of Zero-shot Models via Variance Reduction
von: Zhu, Beier, et al.
Veröffentlicht: (2024) -
Dynamic Multimodal Prototype Learning in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025) -
LoRA of Change: Learning to Generate LoRA for the Editing Instruction from A Single Before-After Image Pair
von: Song, Xue, et al.
Veröffentlicht: (2024) -
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
von: Zhou, Yuan, et al.
Veröffentlicht: (2024)