VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
Fuente:
arXiv
Saved in:
| Main Authors: | Dong, Mingkang, Cai, Hongyi, Li, Jie, Zhou, Sifan, Ren, Bin, Peng, Kunyu, Fu, Yuqian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026)
by: Dong, Mingkang, et al.
Published: (2026)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
by: Ma, Yiwei, et al.
Published: (2025)
by: Ma, Yiwei, et al.
Published: (2025)
Osprey: Pixel Understanding with Visual Instruction Tuning
by: Yuan, Yuqian, et al.
Published: (2023)
by: Yuan, Yuqian, et al.
Published: (2023)
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025)
by: Li, Jie, et al.
Published: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
by: Wang, Dongsheng, et al.
Published: (2024)
by: Wang, Dongsheng, et al.
Published: (2024)
MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
by: Zhang, Renrui, et al.
Published: (2024)
by: Zhang, Renrui, et al.
Published: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
by: He, Junwen, et al.
Published: (2024)
by: He, Junwen, et al.
Published: (2024)
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
by: Wang, Zirui, et al.
Published: (2026)
by: Wang, Zirui, et al.
Published: (2026)
Personalized Visual Instruction Tuning
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Generative Visual Instruction Tuning
by: Hernandez, Jefferson, et al.
Published: (2024)
by: Hernandez, Jefferson, et al.
Published: (2024)
Comparison Visual Instruction Tuning
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
by: Wu, Changti, et al.
Published: (2026)
by: Wu, Changti, et al.
Published: (2026)
VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
by: Deng, Jie, et al.
Published: (2026)
by: Deng, Jie, et al.
Published: (2026)
VTPerception-R1: Enhancing Multimodal Reasoning via Explicit Visual and Textual Perceptual Grounding
by: Ding, Yizhuo, et al.
Published: (2025)
by: Ding, Yizhuo, et al.
Published: (2025)
Multimodal Instruction Tuning with Hybrid State Space Models
by: Zhou, Jianing, et al.
Published: (2024)
by: Zhou, Jianing, et al.
Published: (2024)
ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
by: Liu, Yuqi, et al.
Published: (2025)
by: Liu, Yuqi, et al.
Published: (2025)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
by: Zhang, Zhengbo, et al.
Published: (2026)
by: Zhang, Zhengbo, et al.
Published: (2026)
AutoDebias: Automated Framework for Debiasing Text-to-Image Models
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
TokenPacker: Efficient Visual Projector for Multimodal LLM
by: Li, Wentong, et al.
Published: (2024)
by: Li, Wentong, et al.
Published: (2024)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
VisCRA: A Visual Chain Reasoning Attack for Jailbreaking Multimodal Large Language Models
by: Sima, Bingrui, et al.
Published: (2025)
by: Sima, Bingrui, et al.
Published: (2025)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
by: Ji, Haonian, et al.
Published: (2025)
by: Ji, Haonian, et al.
Published: (2025)
CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking
by: Zhou, Sifan, et al.
Published: (2025)
by: Zhou, Sifan, et al.
Published: (2025)
Visual Instruction Tuning with Chain of Region-of-Interest
by: Chen, Yixin, et al.
Published: (2025)
by: Chen, Yixin, et al.
Published: (2025)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
by: Lei, Zeyu, et al.
Published: (2025)
by: Lei, Zeyu, et al.
Published: (2025)
MergeIT: From Selection to Merging for Efficient Instruction Tuning
by: Cai, Hongyi, et al.
Published: (2025)
by: Cai, Hongyi, et al.
Published: (2025)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
EAGLE: Elevating Geometric Reasoning through LLM-empowered Visual Instruction Tuning
by: Li, Zhihao, et al.
Published: (2024)
by: Li, Zhihao, et al.
Published: (2024)
Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
by: Li, Songze, et al.
Published: (2025)
by: Li, Songze, et al.
Published: (2025)
Multimodal Instruction Tuning with Conditional Mixture of LoRA
by: Shen, Ying, et al.
Published: (2024)
by: Shen, Ying, et al.
Published: (2024)
GeoSense: Internalizing Geometric Necessity Perception for Multimodal Reasoning
by: Liu, Ruiheng, et al.
Published: (2026)
by: Liu, Ruiheng, et al.
Published: (2026)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
by: Zhang, Bin, et al.
Published: (2025)
by: Zhang, Bin, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
UAV-VisLoc: A Large-scale Dataset for UAV Visual Localization
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning
by: Deng, Pei, et al.
Published: (2024)
by: Deng, Pei, et al.
Published: (2024)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
by: Peng, Wujian, et al.
Published: (2024)
by: Peng, Wujian, et al.
Published: (2024)
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning
by: Zeng, Xingchen, et al.
Published: (2024)
by: Zeng, Xingchen, et al.
Published: (2024)
EmoVIT: Revolutionizing Emotion Insights with Visual Instruction Tuning
by: Xie, Hongxia, et al.
Published: (2024)
by: Xie, Hongxia, et al.
Published: (2024)
Similar Items
-
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
by: Dong, Mingkang, et al.
Published: (2026) -
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
by: Ma, Yiwei, et al.
Published: (2025) -
Osprey: Pixel Understanding with Visual Instruction Tuning
by: Yuan, Yuqian, et al.
Published: (2023) -
Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
by: Li, Jie, et al.
Published: (2025) -
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)