VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Guanyu, Yin, Yida, Chai, Wenhao, Tong, Shengbang, Fu, Xingyu, Liu, Zhuang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
UEval: A Benchmark for Unified Multimodal Generation
por: Li, Bo, et al.
Publicado: (2026)
por: Li, Bo, et al.
Publicado: (2026)
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
por: Shahgir, Haz Sameen, et al.
Publicado: (2026)
por: Shahgir, Haz Sameen, et al.
Publicado: (2026)
On the Perception Bottleneck of VLMs for Chart Understanding
por: Liu, Junteng, et al.
Publicado: (2025)
por: Liu, Junteng, et al.
Publicado: (2025)
Reinforced Visual Perception with Tools
por: Zhou, Zetong, et al.
Publicado: (2025)
por: Zhou, Zetong, et al.
Publicado: (2025)
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
por: Shakhadri, Syed Abdul Gaffar, et al.
Publicado: (2025)
por: Shakhadri, Syed Abdul Gaffar, et al.
Publicado: (2025)
Understanding and Rectifying Safety Perception Distortion in VLMs
por: Zou, Xiaohan, et al.
Publicado: (2025)
por: Zou, Xiaohan, et al.
Publicado: (2025)
Teaching Vision-Language Models to Ask: Resolving Ambiguity in Visual Questions
por: Jian, Pu, et al.
Publicado: (2025)
por: Jian, Pu, et al.
Publicado: (2025)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
por: Wang, Wenbin, et al.
Publicado: (2025)
por: Wang, Wenbin, et al.
Publicado: (2025)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
por: Shi, Chufan, et al.
Publicado: (2026)
por: Shi, Chufan, et al.
Publicado: (2026)
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
por: Fu, Xingyu, et al.
Publicado: (2025)
por: Fu, Xingyu, et al.
Publicado: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
por: Dai, Haocheng, et al.
Publicado: (2024)
por: Dai, Haocheng, et al.
Publicado: (2024)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
por: Liu, Peng, et al.
Publicado: (2025)
por: Liu, Peng, et al.
Publicado: (2025)
BabyVision: Visual Reasoning Beyond Language
por: Chen, Liang, et al.
Publicado: (2026)
por: Chen, Liang, et al.
Publicado: (2026)
Smart Eyes for Silent Threats: VLMs and In-Context Learning for THz Imaging
por: Poggi, Nicolas, et al.
Publicado: (2025)
por: Poggi, Nicolas, et al.
Publicado: (2025)
Are VLMs Really Blind
por: Singh, Ayush, et al.
Publicado: (2024)
por: Singh, Ayush, et al.
Publicado: (2024)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
por: Wang, Zhaokai, et al.
Publicado: (2025)
por: Wang, Zhaokai, et al.
Publicado: (2025)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
por: Ye, Junyi, et al.
Publicado: (2024)
por: Ye, Junyi, et al.
Publicado: (2024)
Teaching Text-to-Image Models to Communicate in Dialog
por: Sun, Xiaowen, et al.
Publicado: (2023)
por: Sun, Xiaowen, et al.
Publicado: (2023)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
por: Chen, Liang, et al.
Publicado: (2025)
por: Chen, Liang, et al.
Publicado: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
por: Zhang, Ce, et al.
Publicado: (2025)
por: Zhang, Ce, et al.
Publicado: (2025)
SemVink: Advancing VLMs' Semantic Understanding of Optical Illusions via Visual Global Thinking
por: Li, Sifan, et al.
Publicado: (2025)
por: Li, Sifan, et al.
Publicado: (2025)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
por: Avogaro, Niccolo, et al.
Publicado: (2026)
por: Avogaro, Niccolo, et al.
Publicado: (2026)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
por: Fu, Xingyu, et al.
Publicado: (2023)
por: Fu, Xingyu, et al.
Publicado: (2023)
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs
por: Nasiriany, Soroush, et al.
Publicado: (2024)
por: Nasiriany, Soroush, et al.
Publicado: (2024)
Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs
por: Tong, Shengbang, et al.
Publicado: (2024)
por: Tong, Shengbang, et al.
Publicado: (2024)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
por: Hu, Yushi, et al.
Publicado: (2024)
por: Hu, Yushi, et al.
Publicado: (2024)
Understanding Bias in Large-Scale Visual Datasets
por: Zeng, Boya, et al.
Publicado: (2024)
por: Zeng, Boya, et al.
Publicado: (2024)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
por: Kamoi, Ryo, et al.
Publicado: (2024)
por: Kamoi, Ryo, et al.
Publicado: (2024)
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
por: Zhang, Juntian, et al.
Publicado: (2025)
por: Zhang, Juntian, et al.
Publicado: (2025)
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
por: Zhang, Yanzhe, et al.
Publicado: (2023)
por: Zhang, Yanzhe, et al.
Publicado: (2023)
CIVET: Systematic Evaluation of Understanding in VLMs
por: Rizzoli, Massimo, et al.
Publicado: (2025)
por: Rizzoli, Massimo, et al.
Publicado: (2025)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
por: Mayer, Julius, et al.
Publicado: (2025)
por: Mayer, Julius, et al.
Publicado: (2025)
Can VLMs Recall Factual Associations From Visual References?
por: Ashok, Dhananjay, et al.
Publicado: (2025)
por: Ashok, Dhananjay, et al.
Publicado: (2025)
Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
por: Wang, Dianyi, et al.
Publicado: (2025)
por: Wang, Dianyi, et al.
Publicado: (2025)
Visual In-Context Learning for Large Vision-Language Models
por: Zhou, Yucheng, et al.
Publicado: (2024)
por: Zhou, Yucheng, et al.
Publicado: (2024)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks
por: Jia, Mengzhao, et al.
Publicado: (2024)
por: Jia, Mengzhao, et al.
Publicado: (2024)
[De|Re]constructing VLMs' Reasoning in Counting
por: Alghisi, Simone, et al.
Publicado: (2025)
por: Alghisi, Simone, et al.
Publicado: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
por: Wu, Yin, et al.
Publicado: (2025)
por: Wu, Yin, et al.
Publicado: (2025)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
por: Jiang, Houcheng, et al.
Publicado: (2026)
por: Jiang, Houcheng, et al.
Publicado: (2026)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
por: Qiao, Yuxuan, et al.
Publicado: (2024)
por: Qiao, Yuxuan, et al.
Publicado: (2024)
Ejemplares similares
-
UEval: A Benchmark for Unified Multimodal Generation
por: Li, Bo, et al.
Publicado: (2026) -
VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors
por: Shahgir, Haz Sameen, et al.
Publicado: (2026) -
On the Perception Bottleneck of VLMs for Chart Understanding
por: Liu, Junteng, et al.
Publicado: (2025) -
Reinforced Visual Perception with Tools
por: Zhou, Zetong, et al.
Publicado: (2025) -
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI
por: Shakhadri, Syed Abdul Gaffar, et al.
Publicado: (2025)