Do large language vision models understand 3D shapes?
Fuente:
arXiv
Saved in:
| Main Author: | Eppel, Sagi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods
by: Eppel, Sagi
Published: (2024)
by: Eppel, Sagi
Published: (2024)
Coding the Visual World: From Image to Simulation Using Vision Language Models
by: Eppel, Sagi
Published: (2026)
by: Eppel, Sagi
Published: (2026)
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
by: Eppel, Sagi, et al.
Published: (2025)
by: Eppel, Sagi, et al.
Published: (2025)
Shape and Texture Recognition in Large Vision-Language Models
by: Eppel, Sagi, et al.
Published: (2025)
by: Eppel, Sagi, et al.
Published: (2025)
Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data
by: Eppel, Sagi, et al.
Published: (2024)
by: Eppel, Sagi, et al.
Published: (2024)
ViSTa Dataset: Do vision-language models understand sequential tasks?
by: Wybitul, Evžen, et al.
Published: (2024)
by: Wybitul, Evžen, et al.
Published: (2024)
One-shot recognition of any material anywhere using contrastive learning with physics-based rendering
by: Drehwald, Manuel S., et al.
Published: (2022)
by: Drehwald, Manuel S., et al.
Published: (2022)
LLaVAction: evaluating and training multi-modal large language models for action understanding
by: Qi, Haozhe, et al.
Published: (2025)
by: Qi, Haozhe, et al.
Published: (2025)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
by: Wu, Yiqi, et al.
Published: (2024)
by: Wu, Yiqi, et al.
Published: (2024)
Representation geometry shapes task performance in vision-language modeling for CT enterography
by: Minoccheri, Cristian, et al.
Published: (2026)
by: Minoccheri, Cristian, et al.
Published: (2026)
Hallucination-aware intermediate representation edit in large vision-language models
by: Suo, Wei, et al.
Published: (2026)
by: Suo, Wei, et al.
Published: (2026)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
Visual hallucination detection in large vision-language models via evidential conflict
by: Huang, Tao, et al.
Published: (2025)
by: Huang, Tao, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
An analysis of vision-language models for fabric retrieval
by: Giuliari, Francesco, et al.
Published: (2025)
by: Giuliari, Francesco, et al.
Published: (2025)
On the test-time zero-shot generalization of vision-language models: Do we really need prompt learning?
by: Zanella, Maxime, et al.
Published: (2024)
by: Zanella, Maxime, et al.
Published: (2024)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Self-adaptive vision-language model for 3D segmentation of pulmonary artery and vein
by: Guo, Xiaotong, et al.
Published: (2025)
by: Guo, Xiaotong, et al.
Published: (2025)
Comprehensive language-image pre-training for 3D medical image understanding
by: Wald, Tassilo, et al.
Published: (2025)
by: Wald, Tassilo, et al.
Published: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Visual symbolic mechanisms: Emergent symbol processing in vision language models
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
by: Ranario, Earl, et al.
Published: (2025)
by: Ranario, Earl, et al.
Published: (2025)
Zero-shot large vision-language model prompting for automated bone identification in paleoradiology x-ray archives
by: Dong, Owen, et al.
Published: (2026)
by: Dong, Owen, et al.
Published: (2026)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Functionality understanding and segmentation in 3D scenes
by: Corsetti, Jaime, et al.
Published: (2024)
by: Corsetti, Jaime, et al.
Published: (2024)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
Learning rigid-body simulators over implicit shapes for large-scale scenes and vision
by: Rubanova, Yulia, et al.
Published: (2024)
by: Rubanova, Yulia, et al.
Published: (2024)
Enhancing medical vision-language contrastive learning via inter-matching relation modelling
by: Li, Mingjian, et al.
Published: (2024)
by: Li, Mingjian, et al.
Published: (2024)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Enhancing the vision-language foundation model with key semantic knowledge-emphasized report refinement
by: Huang, Weijian, et al.
Published: (2024)
by: Huang, Weijian, et al.
Published: (2024)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
by: Liu, Xiaohong, et al.
Published: (2024)
by: Liu, Xiaohong, et al.
Published: (2024)
Initialization matters in few-shot adaptation of vision-language models for histopathological image classification
by: Meseguer, Pablo, et al.
Published: (2026)
by: Meseguer, Pablo, et al.
Published: (2026)
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Quantifying the human visual exposome with vision language models
by: Rominger, Christian, et al.
Published: (2026)
by: Rominger, Christian, et al.
Published: (2026)
Near, far: Patch-ordering enhances vision foundation models' scene understanding
by: Pariza, Valentinos, et al.
Published: (2024)
by: Pariza, Valentinos, et al.
Published: (2024)
ShapeFusion: A 3D diffusion model for localized shape editing
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
by: Potamias, Rolandos Alexandros, et al.
Published: (2024)
In-context learning enables multimodal large language models to classify cancer pathology images
by: Ferber, Dyke, et al.
Published: (2024)
by: Ferber, Dyke, et al.
Published: (2024)
Do vision models perceive illusory motion in static images like humans?
by: Rosario, Isabella Elaine, et al.
Published: (2026)
by: Rosario, Isabella Elaine, et al.
Published: (2026)
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Similar Items
-
Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods
by: Eppel, Sagi
Published: (2024) -
Coding the Visual World: From Image to Simulation Using Vision Language Models
by: Eppel, Sagi
Published: (2026) -
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
by: Eppel, Sagi, et al.
Published: (2025) -
Shape and Texture Recognition in Large Vision-Language Models
by: Eppel, Sagi, et al.
Published: (2025) -
Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data
by: Eppel, Sagi, et al.
Published: (2024)