Lever LM: Configuring In-Context Sequence to Lever Large Vision Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Xu, Peng, Yingzhe, Ma, Haoxuan, Xu, Shuo, Zhang, Chi, Han, Yucheng, Zhang, Hanwang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
von: Woo, Sangmin, et al.
Veröffentlicht: (2024)
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
von: Zhang, Pan, et al.
Veröffentlicht: (2024)
von: Zhang, Pan, et al.
Veröffentlicht: (2024)
VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
von: Wang, Zihu, et al.
Veröffentlicht: (2025)
InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024)
VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
von: Shen, Haozhan, et al.
Veröffentlicht: (2025)
von: Shen, Haozhan, et al.
Veröffentlicht: (2025)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
von: Luo, Fuwen, et al.
Veröffentlicht: (2024)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
IdealGPT: Iteratively Decomposing Vision and Language Reasoning via Large Language Models
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
von: You, Haoxuan, et al.
Veröffentlicht: (2023)
HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
von: Wei, Yanbin, et al.
Veröffentlicht: (2026)
Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
von: Jiang, Lei, et al.
Veröffentlicht: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models
von: Lei, Xuanyu, et al.
Veröffentlicht: (2024)
von: Lei, Xuanyu, et al.
Veröffentlicht: (2024)
VLP: A Survey on Vision-Language Pre-training
von: Chen, Feilong, et al.
Veröffentlicht: (2022)
von: Chen, Feilong, et al.
Veröffentlicht: (2022)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
von: Wu, Yuhang, et al.
Veröffentlicht: (2024)
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
von: Lu, Jiaying, et al.
Veröffentlicht: (2023)
Probing and Inducing Combinational Creativity in Vision-Language Models
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
von: Peng, Yongqian, et al.
Veröffentlicht: (2025)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
von: Gao, Peng, et al.
Veröffentlicht: (2021)
von: Gao, Peng, et al.
Veröffentlicht: (2021)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
Intriguing Properties of Large Language and Vision Models
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
von: Lee, Young-Jun, et al.
Veröffentlicht: (2024)
A Unified Hallucination Mitigation Framework for Large Vision-Language Models
von: Chang, Yue, et al.
Veröffentlicht: (2024)
von: Chang, Yue, et al.
Veröffentlicht: (2024)
LADR: Locality-Aware Dynamic Rescue for Efficient Text-to-Image Generation with Diffusion Large Language Models
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
von: Wang, Chenglin, et al.
Veröffentlicht: (2026)
Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2025)
Can Large Vision-Language Models Understand Multimodal Sarcasm?
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
von: Wang, Xinyu, et al.
Veröffentlicht: (2025)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Diving into Mitigating Hallucinations from a Vision Perspective for Large Vision-Language Models
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
von: Wang, Weihang, et al.
Veröffentlicht: (2025)
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
von: Lu, Xudong, et al.
Veröffentlicht: (2024)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
von: Wu, Xueqing, et al.
Veröffentlicht: (2026)
MMLongCite: A Benchmark for Evaluating Fidelity of Long-Context Vision-Language Models
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
von: Zhou, Keyan, et al.
Veröffentlicht: (2025)
Multi-Object Hallucination in Vision-Language Models
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
von: Chen, Xuweiyi, et al.
Veröffentlicht: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
von: Sun, Min Woo, et al.
Veröffentlicht: (2025)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
von: Wu, Daiqing, et al.
Veröffentlicht: (2025)
Trajectory Prediction Meets Large Language Models: A Survey
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
von: Ma, Ruiqi, et al.
Veröffentlicht: (2025)
Kosmos-G: Generating Images in Context with Multimodal Large Language Models
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
von: Pan, Xichen, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024) -
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024) -
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
von: Woo, Sangmin, et al.
Veröffentlicht: (2024) -
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2024) -
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
von: Zhang, Pan, et al.
Veröffentlicht: (2024)