Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lei, Xuanyu, Yang, Zonghan, Chen, Xinrui, Li, Peng, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
von: Gao, Junyuan, et al.
Veröffentlicht: (2025)
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
von: Zhang, Yin, et al.
Veröffentlicht: (2026)
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
von: Huang, Qidong, et al.
Veröffentlicht: (2024)
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models
von: Liu, Jie, et al.
Veröffentlicht: (2024)
von: Liu, Jie, et al.
Veröffentlicht: (2024)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
Coordinated Robustness Evaluation Framework for Vision-Language Models
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
von: Babu, Ashwin Ramesh, et al.
Veröffentlicht: (2025)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
von: Xu, Runsen, et al.
Veröffentlicht: (2025)
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
von: Ju, Tianjie, et al.
Veröffentlicht: (2025)
Otter: A Multi-Modal Model with In-Context Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
Vision-Centric Activation and Coordination for Multimodal Large Language Models
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
von: Wang, Yunnan, et al.
Veröffentlicht: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
From Pixels to Tokens: Revisiting Object Hallucinations in Large Vision-Language Models
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
von: Shang, Yuying, et al.
Veröffentlicht: (2024)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
von: Cheng, Sijie, et al.
Veröffentlicht: (2023)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
A Survey on Hallucination in Large Vision-Language Models
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
von: Liu, Hanchao, et al.
Veröffentlicht: (2024)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
Instruction-Aligned Visual Attention for Mitigating Hallucinations in Large Vision-Language Models
von: Li, Bin, et al.
Veröffentlicht: (2025)
von: Li, Bin, et al.
Veröffentlicht: (2025)
Inference Compute-Optimal Video Vision Language Models
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
von: Wang, Peiqi, et al.
Veröffentlicht: (2025)
Modality-Specialized Synergizers for Interleaved Vision-Language Generalists
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyang, et al.
Veröffentlicht: (2024)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
von: Shao, Zhenwei, et al.
Veröffentlicht: (2025)
von: Shao, Zhenwei, et al.
Veröffentlicht: (2025)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
von: Xu, Shicheng, et al.
Veröffentlicht: (2024)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
von: Zhao, Dachuan, et al.
Veröffentlicht: (2025)
Mitigating Hallucinations in Large Vision-Language Models with Internal Fact-based Contrastive Decoding
von: Wang, Chao, et al.
Veröffentlicht: (2025)
von: Wang, Chao, et al.
Veröffentlicht: (2025)
VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
von: Wang, Xiasi, et al.
Veröffentlicht: (2025)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception
von: Luo, Run, et al.
Veröffentlicht: (2024)
von: Luo, Run, et al.
Veröffentlicht: (2024)
VL-RewardBench: A Challenging Benchmark for Vision-Language Generative Reward Models
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
SAM4MLLM: Enhance Multi-Modal Large Language Model for Referring Expression Segmentation
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
von: Chen, Yi-Chia, et al.
Veröffentlicht: (2024)
Image-Based Geolocation Using Large Vision-Language Models
von: Liu, Yi, et al.
Veröffentlicht: (2024)
von: Liu, Yi, et al.
Veröffentlicht: (2024)
Aligning Modalities in Vision Large Language Models via Preference Fine-tuning
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2024)
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
von: Xu, Yige, et al.
Veröffentlicht: (2026)
von: Xu, Yige, et al.
Veröffentlicht: (2026)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2025) -
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024) -
PM4Bench: Benchmarking Large Vision-Language Models with Parallel Multilingual Multi-Modal Multi-task Corpus
von: Gao, Junyuan, et al.
Veröffentlicht: (2025) -
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment
von: Li, Lei, et al.
Veröffentlicht: (2024) -
MIRL: Mutual Information-Guided Reinforcement Learning for Vision-Language Models
von: Zhang, Yin, et al.
Veröffentlicht: (2026)