Evaluating Graphical Perception Capabilities of Vision Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Poonam, Poonam, Vázquez, Pere-Pau, Ropinski, Timo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Survey on Quality Metrics for Text-to-Image Generation
di: Hartwig, Sebastian, et al.
Pubblicazione: (2024)
di: Hartwig, Sebastian, et al.
Pubblicazione: (2024)
Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
di: Engel, Dominik, et al.
Pubblicazione: (2023)
di: Engel, Dominik, et al.
Pubblicazione: (2023)
Active Learning Inspired ControlNet Guidance for Augmenting Semantic Segmentation Datasets
di: Kniesel, Hannah, et al.
Pubblicazione: (2025)
di: Kniesel, Hannah, et al.
Pubblicazione: (2025)
Unified Semantic Transformer for 3D Scene Understanding
di: Koch, Sebastian, et al.
Pubblicazione: (2025)
di: Koch, Sebastian, et al.
Pubblicazione: (2025)
Unsupervised Semantic Segmentation Through Depth-Guided Feature Correlation and Sampling
di: Sick, Leon, et al.
Pubblicazione: (2023)
di: Sick, Leon, et al.
Pubblicazione: (2023)
DEF-YOLO: Leveraging YOLO for Concealed Weapon Detection in Thermal Imagin
di: Bhardwaj, Divya, et al.
Pubblicazione: (2025)
di: Bhardwaj, Divya, et al.
Pubblicazione: (2025)
StrideNET: Swin Transformer for Terrain Recognition with Dynamic Roughness Extraction
di: Shelare, Maitreya, et al.
Pubblicazione: (2024)
di: Shelare, Maitreya, et al.
Pubblicazione: (2024)
Evaluating Graphical Perception with Multimodal LLMs
di: Nguyen, Rami Huu, et al.
Pubblicazione: (2025)
di: Nguyen, Rami Huu, et al.
Pubblicazione: (2025)
Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images
di: Wolf, Daniel, et al.
Pubblicazione: (2025)
di: Wolf, Daniel, et al.
Pubblicazione: (2025)
Open3DSG: Open-Vocabulary 3D Scene Graphs from Point Clouds with Queryable Objects and Open-Set Relationships
di: Koch, Sebastian, et al.
Pubblicazione: (2024)
di: Koch, Sebastian, et al.
Pubblicazione: (2024)
S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation
di: Sick, Leon, et al.
Pubblicazione: (2025)
di: Sick, Leon, et al.
Pubblicazione: (2025)
CutS3D: Cutting Semantics in 3D for 2D Unsupervised Instance Segmentation
di: Sick, Leon, et al.
Pubblicazione: (2024)
di: Sick, Leon, et al.
Pubblicazione: (2024)
OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
di: Weijler, Lisa, et al.
Pubblicazione: (2025)
di: Weijler, Lisa, et al.
Pubblicazione: (2025)
Attention-Guided Masked Autoencoders For Learning Image Representations
di: Sick, Leon, et al.
Pubblicazione: (2024)
di: Sick, Leon, et al.
Pubblicazione: (2024)
RelationField: Relate Anything in Radiance Fields
di: Koch, Sebastian, et al.
Pubblicazione: (2024)
di: Koch, Sebastian, et al.
Pubblicazione: (2024)
Graphical Perception of Saliency-based Model Explanations
di: Zhao, Yayan, et al.
Pubblicazione: (2024)
di: Zhao, Yayan, et al.
Pubblicazione: (2024)
Unleashing the Capabilities of Large Vision-Language Models for Intelligent Perception of Roadside Infrastructure
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
di: Fu, Luxuan, et al.
Pubblicazione: (2026)
Depth-Wise Representation Development Under Blockwise Self-Supervised Learning for Video Vision Transformers
di: Römer, Jonas, et al.
Pubblicazione: (2026)
di: Römer, Jonas, et al.
Pubblicazione: (2026)
Weakly Supervised Virus Capsid Detection with Image-Level Annotations in Electron Microscopy Images
di: Kniesel, Hannah, et al.
Pubblicazione: (2025)
di: Kniesel, Hannah, et al.
Pubblicazione: (2025)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
di: Guo, Grace, et al.
Pubblicazione: (2024)
di: Guo, Grace, et al.
Pubblicazione: (2024)
Can Vision Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective
di: An, Arctanx, et al.
Pubblicazione: (2026)
di: An, Arctanx, et al.
Pubblicazione: (2026)
AFFMAE: Scalable and Efficient Vision Pretraining for Desktop Graphics Cards
di: Smerkous, David, et al.
Pubblicazione: (2026)
di: Smerkous, David, et al.
Pubblicazione: (2026)
ObjectTransforms for Uncertainty Quantification and Reduction in Vision-Based Perception for Autonomous Vehicles
di: Sahu, Nishad, et al.
Pubblicazione: (2025)
di: Sahu, Nishad, et al.
Pubblicazione: (2025)
Bioimpedance a Diagnostic Tool for Tobacco Induced Oral Lesions: a Mixed Model cross-sectional study
di: Gupta, Vaibhav, et al.
Pubblicazione: (2024)
di: Gupta, Vaibhav, et al.
Pubblicazione: (2024)
Evaluating the Explainability of Vision Transformers in Medical Imaging
di: Barekatain, Leili, et al.
Pubblicazione: (2025)
di: Barekatain, Leili, et al.
Pubblicazione: (2025)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
di: Danier, Duolikun, et al.
Pubblicazione: (2024)
di: Danier, Duolikun, et al.
Pubblicazione: (2024)
Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
di: He, Guangzhao, et al.
Pubblicazione: (2026)
di: He, Guangzhao, et al.
Pubblicazione: (2026)
StarVector: Generating Scalable Vector Graphics Code from Images and Text
di: Rodriguez, Juan A., et al.
Pubblicazione: (2023)
di: Rodriguez, Juan A., et al.
Pubblicazione: (2023)
Sim2Real Transfer for Vision-Based Grasp Verification
di: Amargant, Pau, et al.
Pubblicazione: (2025)
di: Amargant, Pau, et al.
Pubblicazione: (2025)
Protego: Detecting Adversarial Examples for Vision Transformers via Intrinsic Capabilities
di: Wu, Jialin, et al.
Pubblicazione: (2025)
di: Wu, Jialin, et al.
Pubblicazione: (2025)
Bridging the Perception Gap in Image Super-Resolution Evaluation
di: Su, Shaolin, et al.
Pubblicazione: (2025)
di: Su, Shaolin, et al.
Pubblicazione: (2025)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
di: Ganj, Ashkan, et al.
Pubblicazione: (2025)
di: Ganj, Ashkan, et al.
Pubblicazione: (2025)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
di: Li, Yu, et al.
Pubblicazione: (2024)
di: Li, Yu, et al.
Pubblicazione: (2024)
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
di: Cheng, Sijie, et al.
Pubblicazione: (2023)
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
di: Zheng, Xiangxi, et al.
Pubblicazione: (2025)
di: Zheng, Xiangxi, et al.
Pubblicazione: (2025)
Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring
di: Onsu, Murat Arda, et al.
Pubblicazione: (2025)
di: Onsu, Murat Arda, et al.
Pubblicazione: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
di: Zhang, Enming, et al.
Pubblicazione: (2025)
di: Zhang, Enming, et al.
Pubblicazione: (2025)
Textured Mesh Saliency: Bridging Geometry and Texture for Human Perception in 3D Graphics
di: Zhang, Kaiwei, et al.
Pubblicazione: (2024)
di: Zhang, Kaiwei, et al.
Pubblicazione: (2024)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
di: Huang, Ian, et al.
Pubblicazione: (2024)
di: Huang, Ian, et al.
Pubblicazione: (2024)
CreatiDesign: A Unified Multi-Conditional Diffusion Transformer for Creative Graphic Design
di: Zhang, Hui, et al.
Pubblicazione: (2025)
di: Zhang, Hui, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Survey on Quality Metrics for Text-to-Image Generation
di: Hartwig, Sebastian, et al.
Pubblicazione: (2024) -
Leveraging Self-Supervised Vision Transformers for Segmentation-based Transfer Function Design
di: Engel, Dominik, et al.
Pubblicazione: (2023) -
Active Learning Inspired ControlNet Guidance for Augmenting Semantic Segmentation Datasets
di: Kniesel, Hannah, et al.
Pubblicazione: (2025) -
Unified Semantic Transformer for 3D Scene Understanding
di: Koch, Sebastian, et al.
Pubblicazione: (2025) -
Unsupervised Semantic Segmentation Through Depth-Guided Feature Correlation and Sampling
di: Sick, Leon, et al.
Pubblicazione: (2023)