AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ganj, Ashkan, Zhao, Yiqin, Guo, Tian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors
by: Ganj, Ashkan, et al.
Published: (2024)
by: Ganj, Ashkan, et al.
Published: (2024)
Can Foundation Models Revolutionize Mobile AR Sparse Sensing?
by: Zhao, Yiqin, et al.
Published: (2025)
by: Zhao, Yiqin, et al.
Published: (2025)
CleAR: Robust Context-Guided Generative Lighting Estimation for Mobile Augmented Reality
by: Zhao, Yiqin, et al.
Published: (2024)
by: Zhao, Yiqin, et al.
Published: (2024)
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
by: Wang, Yiqin, et al.
Published: (2024)
by: Wang, Yiqin, et al.
Published: (2024)
Bridging Classical and Modern Computer Vision: PerceptiveNet for Tree Crown Semantic Segmentation
by: Voulgaris, Georgios
Published: (2025)
by: Voulgaris, Georgios
Published: (2025)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
Not There Yet: Evaluating Vision Language Models in Simulating the Visual Perception of People with Low Vision
by: Natalie, Rosiana, et al.
Published: (2025)
by: Natalie, Rosiana, et al.
Published: (2025)
Same or Not? Enhancing Visual Perception in Vision-Language Models
by: Marsili, Damiano, et al.
Published: (2025)
by: Marsili, Damiano, et al.
Published: (2025)
Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
by: Dong, Shiyin, et al.
Published: (2024)
by: Dong, Shiyin, et al.
Published: (2024)
Learning Visual Prompts for Guiding the Attention of Vision Transformers
by: Rezaei, Razieh, et al.
Published: (2024)
by: Rezaei, Razieh, et al.
Published: (2024)
TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
by: Zhou, Wenhao, et al.
Published: (2025)
by: Zhou, Wenhao, et al.
Published: (2025)
Hidden in Plain Sight: Evaluating Abstract Shape Recognition in Vision-Language Models
by: Hemmat, Arshia, et al.
Published: (2024)
by: Hemmat, Arshia, et al.
Published: (2024)
Selective Vision is the Challenge for Visual Reasoning: A Benchmark for Visual Argument Understanding
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Bridging the Perception Gap in Image Super-Resolution Evaluation
by: Su, Shaolin, et al.
Published: (2025)
by: Su, Shaolin, et al.
Published: (2025)
CARLA-GeAR: a Dataset Generator for a Systematic Evaluation of Adversarial Robustness of Vision Models
by: Nesti, Federico, et al.
Published: (2022)
by: Nesti, Federico, et al.
Published: (2022)
Unveiling and Bridging the Functional Perception Gap in MLLMs: Atomic Visual Alignment and Hierarchical Evaluation via PET-Bench
by: Ye, Zanting, et al.
Published: (2026)
by: Ye, Zanting, et al.
Published: (2026)
DepthCues: Evaluating Monocular Depth Perception in Large Vision Models
by: Danier, Duolikun, et al.
Published: (2024)
by: Danier, Duolikun, et al.
Published: (2024)
Evaluating Graphical Perception Capabilities of Vision Transformers
by: Poonam, Poonam, et al.
Published: (2026)
by: Poonam, Poonam, et al.
Published: (2026)
Systematic Literature Review on Vehicular Collaborative Perception -- A Computer Vision Perspective
by: Wan, Lei, et al.
Published: (2025)
by: Wan, Lei, et al.
Published: (2025)
Bridging Visual Representation and Reinforcement Learning from Verifiable Rewards in Large Vision-Language Models
by: Han, Yuhang, et al.
Published: (2026)
by: Han, Yuhang, et al.
Published: (2026)
DepthLM: Metric Depth From Vision Language Models
by: Cai, Zhipeng, et al.
Published: (2025)
by: Cai, Zhipeng, et al.
Published: (2025)
A Comprehensive Safety Metric to Evaluate Perception in Autonomous Systems
by: Volk, Georg, et al.
Published: (2025)
by: Volk, Georg, et al.
Published: (2025)
SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
by: Wang, Junke, et al.
Published: (2025)
by: Wang, Junke, et al.
Published: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
by: Dai, Haocheng, et al.
Published: (2024)
by: Dai, Haocheng, et al.
Published: (2024)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
PolarVLM: Bridging the Semantic-Physical Gap in Vision-Language Models
by: Li, Yuliang, et al.
Published: (2026)
by: Li, Yuliang, et al.
Published: (2026)
Learnable Sparsity for Vision Generative Models
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
by: Duan, Yuchen, et al.
Published: (2024)
by: Duan, Yuchen, et al.
Published: (2024)
Do Vision-Language Foundational models show Robust Visual Perception?
by: Chandhok, Shivam, et al.
Published: (2024)
by: Chandhok, Shivam, et al.
Published: (2024)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
by: Guo, Grace, et al.
Published: (2024)
by: Guo, Grace, et al.
Published: (2024)
Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models
by: Zhan, Yufei, et al.
Published: (2024)
by: Zhan, Yufei, et al.
Published: (2024)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
by: Zhao, Shihao, et al.
Published: (2024)
by: Zhao, Shihao, et al.
Published: (2024)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
by: Lalai, Harsh Nishant, et al.
Published: (2026)
by: Lalai, Harsh Nishant, et al.
Published: (2026)
Evaluating Self-Correcting Vision Agents Through Quantitative and Qualitative Metrics
by: Dixit, Aradhya
Published: (2026)
by: Dixit, Aradhya
Published: (2026)
Active Visual Perception: Opportunities and Challenges
by: Li, Yian, et al.
Published: (2025)
by: Li, Yian, et al.
Published: (2025)
Improving Vision-language Models with Perception-centric Process Reward Models
by: Min, Yingqian, et al.
Published: (2026)
by: Min, Yingqian, et al.
Published: (2026)
MetaRank: Task-Aware Metric Selection for Model Transferability Estimation
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
Position: Evaluation of Visual Processing Should Be Human-Centered, Not Metric-Centered
by: Hu, Jinfan, et al.
Published: (2026)
by: Hu, Jinfan, et al.
Published: (2026)
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
Impact of Target and Tool Visualization on Depth Perception and Usability in Optical See-Through AR
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
Similar Items
-
HybridDepth: Robust Metric Depth Fusion by Leveraging Depth from Focus and Single-Image Priors
by: Ganj, Ashkan, et al.
Published: (2024) -
Can Foundation Models Revolutionize Mobile AR Sparse Sensing?
by: Zhao, Yiqin, et al.
Published: (2025) -
CleAR: Robust Context-Guided Generative Lighting Estimation for Mobile Augmented Reality
by: Zhao, Yiqin, et al.
Published: (2024) -
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control
by: Wang, Yiqin, et al.
Published: (2024) -
Bridging Classical and Modern Computer Vision: PerceptiveNet for Tree Crown Semantic Segmentation
by: Voulgaris, Georgios
Published: (2025)