VLLaVO: Mitigating Visual Gap through LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Shuhao, Zhang, Yulong, Jiang, Weisen, Lu, Jiangang, Zhang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BYOM: Building Your Own Multi-Task Model For Free
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
von: Jiang, Weisen, et al.
Veröffentlicht: (2023)
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
von: Han, Songhao, et al.
Veröffentlicht: (2023)
von: Han, Songhao, et al.
Veröffentlicht: (2023)
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
von: Li, Kaican, et al.
Veröffentlicht: (2025)
von: Li, Kaican, et al.
Veröffentlicht: (2025)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
Mitigating the Modality Gap: Few-Shot Out-of-Distribution Detection with Multi-modal Prototypes and Image Bias Estimation
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
von: Wang, Yimu, et al.
Veröffentlicht: (2025)
EMMA: Efficient Visual Alignment in Multi-Modal LLMs
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2024)
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
von: Deng, Naihao, et al.
Veröffentlicht: (2024)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
von: Eslami, Sedigheh, et al.
Veröffentlicht: (2024)
MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimization
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
von: Chaubey, Ashutosh, et al.
Veröffentlicht: (2026)
Composition-Grounded Data Synthesis for Visual Reasoning
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2023)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
von: You, Zebin, et al.
Veröffentlicht: (2025)
von: You, Zebin, et al.
Veröffentlicht: (2025)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
von: Chang, Kai-Po, et al.
Veröffentlicht: (2025)
von: Chang, Kai-Po, et al.
Veröffentlicht: (2025)
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
von: Meshram, Pragati Shuddhodhan, et al.
Veröffentlicht: (2024)
Find The Gap: Knowledge Base Reasoning For Visual Question Answering
von: Barezi, Elham J., et al.
Veröffentlicht: (2024)
von: Barezi, Elham J., et al.
Veröffentlicht: (2024)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
MoPD: Mixture-of-Prompts Distillation for Vision-Language Models
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Attribute Diversity Determines the Systematicity Gap in VQA
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
Rethinking Guidance Information to Utilize Unlabeled Samples:A Label Encoding Perspective
von: Zhang, Yulong, et al.
Veröffentlicht: (2024)
von: Zhang, Yulong, et al.
Veröffentlicht: (2024)
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Systematic Reward Gap Optimization for Mitigating VLM Hallucinations
von: He, Lehan, et al.
Veröffentlicht: (2024)
von: He, Lehan, et al.
Veröffentlicht: (2024)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Breaking through the learning plateaus of in-context learning in Transformer
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
Reefknot: A Comprehensive Benchmark for Relation Hallucination Evaluation, Analysis and Mitigation in Multimodal Large Language Models
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
von: Zheng, Kening, et al.
Veröffentlicht: (2024)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
von: Shenoy, Ashish, et al.
Veröffentlicht: (2024)
von: Shenoy, Ashish, et al.
Veröffentlicht: (2024)
Prophet: Prompting Large Language Models with Complementary Answer Heuristics for Knowledge-based Visual Question Answering
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
von: Yu, Zhou, et al.
Veröffentlicht: (2023)
LDP: Generalizing to Multilingual Visual Information Extraction by Language Decoupled Pretraining
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
von: Shen, Huawen, et al.
Veröffentlicht: (2024)
Gla-AI4BioMed at RRG24: Visual Instruction-tuned Adaptation for Radiology Report Generation
von: Zhang, Xi, et al.
Veröffentlicht: (2024)
von: Zhang, Xi, et al.
Veröffentlicht: (2024)
Evaluating the Correctness of Inference Patterns Used by LLMs for Judgment
von: Chen, Lu, et al.
Veröffentlicht: (2024)
von: Chen, Lu, et al.
Veröffentlicht: (2024)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
von: Gado, Mohamed, et al.
Veröffentlicht: (2025)
MouSi: Poly-Visual-Expert Vision-Language Models
von: Fan, Xiaoran, et al.
Veröffentlicht: (2024)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2024)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
von: Wang, Chuhan, et al.
Veröffentlicht: (2026)
The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
von: Jung, Hoin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BYOM: Building Your Own Multi-Task Model For Free
von: Jiang, Weisen, et al.
Veröffentlicht: (2023) -
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion
von: Chen, Peiyuan, et al.
Veröffentlicht: (2024) -
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025) -
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
von: Han, Songhao, et al.
Veröffentlicht: (2023) -
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
von: Li, Kaican, et al.
Veröffentlicht: (2025)