Shape and Texture Recognition in Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Eppel, Sagi, Bismut, Mor, Faktor-Strugatski, Alona |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
by: Eppel, Sagi, et al.
Published: (2025)
by: Eppel, Sagi, et al.
Published: (2025)
Coding the Visual World: From Image to Simulation Using Vision Language Models
by: Eppel, Sagi
Published: (2026)
by: Eppel, Sagi
Published: (2026)
Do large language vision models understand 3D shapes?
by: Eppel, Sagi
Published: (2024)
by: Eppel, Sagi
Published: (2024)
Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods
by: Eppel, Sagi
Published: (2024)
by: Eppel, Sagi
Published: (2024)
Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data
by: Eppel, Sagi, et al.
Published: (2024)
by: Eppel, Sagi, et al.
Published: (2024)
One-shot recognition of any material anywhere using contrastive learning with physics-based rendering
by: Drehwald, Manuel S., et al.
Published: (2022)
by: Drehwald, Manuel S., et al.
Published: (2022)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
by: Li, Zhecheng, et al.
Published: (2025)
by: Li, Zhecheng, et al.
Published: (2025)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Contextual Emotion Recognition using Large Vision Language Models
by: Etesam, Yasaman, et al.
Published: (2024)
by: Etesam, Yasaman, et al.
Published: (2024)
IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models
by: Shi, Liang, et al.
Published: (2026)
by: Shi, Liang, et al.
Published: (2026)
Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
by: Li, Ling, et al.
Published: (2025)
by: Li, Ling, et al.
Published: (2025)
Leveraging Vision-Language Large Models for Interpretable Video Action Recognition with Semantic Tokenization
by: Peng, Jingwei, et al.
Published: (2025)
by: Peng, Jingwei, et al.
Published: (2025)
Compound Expression Recognition via Large Vision-Language Models
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
by: Ranasinghe, Yasiru, et al.
Published: (2025)
by: Ranasinghe, Yasiru, et al.
Published: (2025)
Open-Set Recognition in the Age of Vision-Language Models
by: Miller, Dimity, et al.
Published: (2024)
by: Miller, Dimity, et al.
Published: (2024)
ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling
by: Zhang, Shuyuan, et al.
Published: (2025)
by: Zhang, Shuyuan, et al.
Published: (2025)
On the Influence of Shape, Texture and Color for Learning Semantic Segmentation
by: Mütze, Annika, et al.
Published: (2024)
by: Mütze, Annika, et al.
Published: (2024)
Evaluating Vision-Language Models for Emotion Recognition
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Make-A-Texture: Fast Shape-Aware Texture Generation in 3 Seconds
by: Xiang, Xiaoyu, et al.
Published: (2024)
by: Xiang, Xiaoyu, et al.
Published: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Toward Semantic-Agnostic and Shape-Aware Vision-Language Segmentation Models
by: Seutin, Corentin, et al.
Published: (2026)
by: Seutin, Corentin, et al.
Published: (2026)
LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition
by: You, Jinghan, et al.
Published: (2025)
by: You, Jinghan, et al.
Published: (2025)
Personalized Large Vision-Language Models
by: Pham, Chau, et al.
Published: (2024)
by: Pham, Chau, et al.
Published: (2024)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
by: Kim, Jeonghwan, et al.
Published: (2024)
by: Kim, Jeonghwan, et al.
Published: (2024)
Benchmarking Large Language Models for Handwritten Text Recognition
by: Crosilla, Giorgia, et al.
Published: (2025)
by: Crosilla, Giorgia, et al.
Published: (2025)
Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
by: Luan, Tianyu, et al.
Published: (2025)
by: Luan, Tianyu, et al.
Published: (2025)
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
by: Saha, Rohan, et al.
Published: (2024)
by: Saha, Rohan, et al.
Published: (2024)
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
ROSA: Reconstructing Object Shape and Appearance Textures by Adaptive Detail Transfer
by: Kaltheuner, Julian, et al.
Published: (2025)
by: Kaltheuner, Julian, et al.
Published: (2025)
Towards Texture- And Shape-Independent 3D Keypoint Estimation in Birds
by: Schmuker, Valentin, et al.
Published: (2025)
by: Schmuker, Valentin, et al.
Published: (2025)
Phantom of Latent for Large Language and Vision Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
Texture- and Shape-based Adversarial Attacks for Overhead Image Vehicle Detection
by: Yeghiazaryan, Mikael, et al.
Published: (2024)
by: Yeghiazaryan, Mikael, et al.
Published: (2024)
CTGAN: Semantic-guided Conditional Texture Generator for 3D Shapes
by: Pan, Yi-Ting, et al.
Published: (2024)
by: Pan, Yi-Ting, et al.
Published: (2024)
ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
by: Li, Zhaoyang, et al.
Published: (2025)
by: Li, Zhaoyang, et al.
Published: (2025)
Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments
by: Nagaonkar, Sankalp, et al.
Published: (2025)
by: Nagaonkar, Sankalp, et al.
Published: (2025)
Layout-Independent License Plate Recognition via Integrated Vision and Language Models
by: Shabaninia, Elham, et al.
Published: (2025)
by: Shabaninia, Elham, et al.
Published: (2025)
TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
by: Zhang, Junyuan, et al.
Published: (2025)
by: Zhang, Junyuan, et al.
Published: (2025)
An Application-Agnostic Automatic Target Recognition System Using Vision Language Models
by: Palladino, Anthony, et al.
Published: (2024)
by: Palladino, Anthony, et al.
Published: (2024)
AI-based Wearable Vision Assistance System for the Visually Impaired: Integrating Real-Time Object Recognition and Contextual Understanding Using Large Vision-Language Models
by: Baig, Mirza Samad Ahmed, et al.
Published: (2024)
by: Baig, Mirza Samad Ahmed, et al.
Published: (2024)
Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models
by: Wang, Haoming, et al.
Published: (2026)
by: Wang, Haoming, et al.
Published: (2026)
Similar Items
-
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
by: Eppel, Sagi, et al.
Published: (2025) -
Coding the Visual World: From Image to Simulation Using Vision Language Models
by: Eppel, Sagi
Published: (2026) -
Do large language vision models understand 3D shapes?
by: Eppel, Sagi
Published: (2024) -
Vastextures: Vast repository of textures and PBR materials extracted from real-world images using unsupervised methods
by: Eppel, Sagi
Published: (2024) -
Learning Zero-Shot Material States Segmentation, by Implanting Natural Image Patterns in Synthetic Data
by: Eppel, Sagi, et al.
Published: (2024)