Advancing vision-language models in front-end development via data synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Tong, Liu, Yashu, Ye, Jieping, Li, Tianyi, Wang, Chao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
von: Aman, Tabinda, et al.
Veröffentlicht: (2025)
von: Aman, Tabinda, et al.
Veröffentlicht: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
Cross-modal linkage risk in clinical vision-language models
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026)
PathAlign: A vision-language model for whole slide images in histopathology
von: Ahmed, Faruk, et al.
Veröffentlicht: (2024)
von: Ahmed, Faruk, et al.
Veröffentlicht: (2024)
Fostering Video Reasoning via Next-Event Prediction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
von: Padlewski, Piotr, et al.
Veröffentlicht: (2024)
von: Padlewski, Piotr, et al.
Veröffentlicht: (2024)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
von: Rädsch, Tim, et al.
Veröffentlicht: (2025)
von: Rädsch, Tim, et al.
Veröffentlicht: (2025)
From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaofeng, et al.
Veröffentlicht: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
von: Albadarneh, Israa A., et al.
Veröffentlicht: (2025)
Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
von: Wang, Guoxin, et al.
Veröffentlicht: (2025)
von: Wang, Guoxin, et al.
Veröffentlicht: (2025)
Are vision language models robust to uncertain inputs?
von: Wang, Xi, et al.
Veröffentlicht: (2025)
von: Wang, Xi, et al.
Veröffentlicht: (2025)
Multimodal Evaluation of Russian-language Architectures
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
von: Chervyakov, Artem, et al.
Veröffentlicht: (2025)
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
von: Zhao, Haozhe, et al.
Veröffentlicht: (2023)
Quantifying the human visual exposome with vision language models
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
von: Rominger, Christian, et al.
Veröffentlicht: (2026)
What matters when building vision-language models?
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
von: Fang, Ye, et al.
Veröffentlicht: (2024)
von: Fang, Ye, et al.
Veröffentlicht: (2024)
Leveraging Open Knowledge for Advancing Task Expertise in Large Language Models
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
von: Yang, Yuncheng, et al.
Veröffentlicht: (2024)
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
von: Pan, Zhiyu, et al.
Veröffentlicht: (2026)
EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training
von: Du, Yiyang, et al.
Veröffentlicht: (2026)
von: Du, Yiyang, et al.
Veröffentlicht: (2026)
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
von: Khan, Shaharukh, et al.
Veröffentlicht: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Thinker: A vision-language foundation model for embodied intelligence
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
von: Pan, Baiyu, et al.
Veröffentlicht: (2026)
Breaking Down and Building Up: Mixture of Skill-Based Vision-and-Language Navigation Agents
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
von: Ma, Tianyi, et al.
Veröffentlicht: (2025)
VLMQ: Token Saliency-Driven Post-Training Quantization for Vision-language Models
von: Xue, Yufei, et al.
Veröffentlicht: (2025)
von: Xue, Yufei, et al.
Veröffentlicht: (2025)
RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
von: Niu, Tianyi, et al.
Veröffentlicht: (2025)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Huang, Wenxuan, et al.
Veröffentlicht: (2024)
Building and better understanding vision-language models: insights and future directions
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
Hallucination-aware intermediate representation edit in large vision-language models
von: Suo, Wei, et al.
Veröffentlicht: (2026)
von: Suo, Wei, et al.
Veröffentlicht: (2026)
Generalizing vision-language models to novel domains: A comprehensive survey
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Q&A Prompts: Discovering Rich Visual Clues through Mining Question-Answer Prompts for VQA requiring Diverse World Knowledge
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
BrainChat: Decoding Semantic Information from fMRI using Vision-language Pretrained Models
von: Huang, Wanaiu
Veröffentlicht: (2024)
von: Huang, Wanaiu
Veröffentlicht: (2024)
Automatic benchmarking of large multimodal models via iterative experiment programming
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
von: Conti, Alessandro, et al.
Veröffentlicht: (2024)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
von: Englebert, Alexandre, et al.
Veröffentlicht: (2024)
von: Englebert, Alexandre, et al.
Veröffentlicht: (2024)
HA-FGOVD: Highlighting Fine-grained Attributes via Explicit Linear Composition for Open-Vocabulary Object Detection
von: Ma, Yuqi, et al.
Veröffentlicht: (2024)
von: Ma, Yuqi, et al.
Veröffentlicht: (2024)
A benchmark multimodal oro-dental dataset for large vision-language models
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
Representation geometry shapes task performance in vision-language modeling for CT enterography
von: Minoccheri, Cristian, et al.
Veröffentlicht: (2026)
von: Minoccheri, Cristian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Owls are wise and foxes are unfaithful: Uncovering animal stereotypes in vision-language models
von: Aman, Tabinda, et al.
Veröffentlicht: (2025) -
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025) -
Cross-modal linkage risk in clinical vision-language models
von: Arasteh, Soroosh Tayebi, et al.
Veröffentlicht: (2026) -
PathAlign: A vision-language model for whole slide images in histopathology
von: Ahmed, Faruk, et al.
Veröffentlicht: (2024) -
Fostering Video Reasoning via Next-Event Prediction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)