MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Jihao, Huang, Xin, Zheng, Jinliang, Liu, Boxiao, Wang, Jia, Yoshie, Osamu, Liu, Yu, Li, Hongsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Vision-Language Model with Unmasked Token Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
von: Wang, Yuhao, et al.
Veröffentlicht: (2024)
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
von: Jia, Yiming, et al.
Veröffentlicht: (2025)
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
von: Liu, Jianyu, et al.
Veröffentlicht: (2025)
von: Liu, Jianyu, et al.
Veröffentlicht: (2025)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2024)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
von: Hu, Hexiang, et al.
Veröffentlicht: (2024)
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
von: Liu, Zikang, et al.
Veröffentlicht: (2025)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
von: Yu, Weihao, et al.
Veröffentlicht: (2023)
Learning to Instruct for Visual Instruction Tuning
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhou, Zhihan, et al.
Veröffentlicht: (2025)
Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models
von: Li, Yifan, et al.
Veröffentlicht: (2024)
von: Li, Yifan, et al.
Veröffentlicht: (2024)
Instruction-Guided Visual Masking
von: Zheng, Jinliang, et al.
Veröffentlicht: (2024)
von: Zheng, Jinliang, et al.
Veröffentlicht: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
von: Yuan, Qianhao, et al.
Veröffentlicht: (2025)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
von: Yu, Weihao, et al.
Veröffentlicht: (2024)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
von: Tian, Changyao, et al.
Veröffentlicht: (2024)
FlagEvalMM: A Flexible Framework for Comprehensive Multimodal Model Evaluation
von: He, Zheqi, et al.
Veröffentlicht: (2025)
von: He, Zheqi, et al.
Veröffentlicht: (2025)
OmniGen2: Towards Instruction-Aligned Multimodal Generation
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
von: Wu, Chenyuan, et al.
Veröffentlicht: (2025)
DecisionNCE: Embodied Multimodal Representations via Implicit Preference Learning
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
von: Li, Jianxiong, et al.
Veröffentlicht: (2024)
Reinforcing Multimodal Reasoning Against Visual Degradation
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
von: Jin, Yiqiao, et al.
Veröffentlicht: (2024)
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
von: Xing, Bohao, et al.
Veröffentlicht: (2025)
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
von: Li, Shilong, et al.
Veröffentlicht: (2025)
von: Li, Shilong, et al.
Veröffentlicht: (2025)
UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
von: Ji, Yifan, et al.
Veröffentlicht: (2026)
Sparkles: Unlocking Chats Across Multiple Images for Multimodal Instruction-Following Models
von: Huang, Yupan, et al.
Veröffentlicht: (2023)
von: Huang, Yupan, et al.
Veröffentlicht: (2023)
OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
von: Liu, Yuliang, et al.
Veröffentlicht: (2023)
InstructSeg: Unifying Instructed Visual Segmentation with Multi-modal Large Language Models
von: Wei, Cong, et al.
Veröffentlicht: (2024)
von: Wei, Cong, et al.
Veröffentlicht: (2024)
OmniParser V2: Structured-Points-of-Thought for Unified Visual Text Parsing and Its Generality to Multimodal Large Language Models
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
von: Yu, Wenwen, et al.
Veröffentlicht: (2025)
MAGIC: Multimodal Alignment & Grounding-aware Instruction Coreset for Vision-Language Models
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
von: Biswas, Shristi Das, et al.
Veröffentlicht: (2026)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
Mask What Matters: Mitigating Object Hallucinations in Multimodal Large Language Models with Object-Aligned Visual Contrastive Decoding
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
von: Chen, Boqi, et al.
Veröffentlicht: (2026)
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models
von: Xue, Le, et al.
Veröffentlicht: (2024)
von: Xue, Le, et al.
Veröffentlicht: (2024)
InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
von: Wan, Zifu, et al.
Veröffentlicht: (2025)
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
InstructEdit: Instruction-based Knowledge Editing for Large Language Models
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
von: Zhang, Ningyu, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
von: Chen, Feilong, et al.
Veröffentlicht: (2025)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing Vision-Language Model with Unmasked Token Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024) -
GLID: Pre-training a Generalist Encoder-Decoder Vision Model
von: Liu, Jihao, et al.
Veröffentlicht: (2024) -
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
von: Wang, Yuhao, et al.
Veröffentlicht: (2024) -
VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
von: Jia, Yiming, et al.
Veröffentlicht: (2025) -
DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
von: Liu, Jianyu, et al.
Veröffentlicht: (2025)