Understanding the Dependence of Perception Model Competency on Regions in an Image
Fuente:
arXiv
Saved in:
| Main Authors: | Pohland, Sara, Tomlin, Claire |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Explaining Low Perception Model Competency with High-Competency Counterfactuals
by: Pohland, Sara, et al.
Published: (2025)
by: Pohland, Sara, et al.
Published: (2025)
Competency-Aware Planning for Probabilistically Safe Navigation Under Perception Uncertainty
by: Pohland, Sara, et al.
Published: (2024)
by: Pohland, Sara, et al.
Published: (2024)
PaRCE: Probabilistic and Reconstruction-based Competency Estimation for CNN-based Image Classification
by: Pohland, Sara, et al.
Published: (2024)
by: Pohland, Sara, et al.
Published: (2024)
A Deep Learning Approach for Augmenting Perceptional Understanding of Histopathology Images
by: Hu, Xiaoqian
Published: (2025)
by: Hu, Xiaoqian
Published: (2025)
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
by: Fang, Tianchen, et al.
Published: (2025)
by: Fang, Tianchen, et al.
Published: (2025)
Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
by: Theodoridis, Nikos, et al.
Published: (2025)
by: Theodoridis, Nikos, et al.
Published: (2025)
Focus, Don't Prune: Identifying Instruction-Relevant Regions for Information-Rich Image Understanding
by: Kwon, Mincheol, et al.
Published: (2026)
by: Kwon, Mincheol, et al.
Published: (2026)
Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning
by: Zhao, Bingchen, et al.
Published: (2024)
by: Zhao, Bingchen, et al.
Published: (2024)
Understanding Graphical Perception in Data Visualization through Zero-shot Prompting of Vision-Language Models
by: Guo, Grace, et al.
Published: (2024)
by: Guo, Grace, et al.
Published: (2024)
PerspectiveNet: Multi-View Perception for Dynamic Scene Understanding
by: Nguyen, Vinh
Published: (2024)
by: Nguyen, Vinh
Published: (2024)
Region-Level Context-Aware Multimodal Understanding
by: Wei, Hongliang, et al.
Published: (2025)
by: Wei, Hongliang, et al.
Published: (2025)
Region-Aware CAM: High-Resolution Weakly-Supervised Defect Segmentation via Salient Region Perception
by: Dong, Hang-Cheng, et al.
Published: (2025)
by: Dong, Hang-Cheng, et al.
Published: (2025)
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
by: Zhang, Zhiqiu, et al.
Published: (2025)
by: Zhang, Zhiqiu, et al.
Published: (2025)
Unsupervised Region-Based Image Editing of Denoising Diffusion Models
by: Li, Zixiang, et al.
Published: (2024)
by: Li, Zixiang, et al.
Published: (2024)
Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
by: Xu, Zhiyang, et al.
Published: (2025)
by: Xu, Zhiyang, et al.
Published: (2025)
GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
by: Ye, Xianhang, et al.
Published: (2025)
by: Ye, Xianhang, et al.
Published: (2025)
Abductive Ego-View Accident Video Understanding for Safe Driving Perception
by: Fang, Jianwu, et al.
Published: (2024)
by: Fang, Jianwu, et al.
Published: (2024)
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
by: Zheng, Haoyu, et al.
Published: (2026)
by: Zheng, Haoyu, et al.
Published: (2026)
Towards Understanding Graphical Perception in Large Multimodal Models
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
Large Language Models Can Understanding Depth from Monocular Images
by: Xia, Zhongyi, et al.
Published: (2024)
by: Xia, Zhongyi, et al.
Published: (2024)
Prompt Guidance and Human Proximal Perception for HOT Prediction with Regional Joint Loss
by: Wang, Yuxiao, et al.
Published: (2025)
by: Wang, Yuxiao, et al.
Published: (2025)
Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection
by: Yakun, Cui, et al.
Published: (2025)
by: Yakun, Cui, et al.
Published: (2025)
Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models
by: Chen, Yule, et al.
Published: (2025)
by: Chen, Yule, et al.
Published: (2025)
RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding
by: Yang, Jihan, et al.
Published: (2023)
by: Yang, Jihan, et al.
Published: (2023)
Hyperspectral Imaging-Based Perception in Autonomous Driving Scenarios: Benchmarking Baseline Semantic Segmentation Models
by: Shah, Imad Ali, et al.
Published: (2024)
by: Shah, Imad Ali, et al.
Published: (2024)
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
by: Guo, Mingning, et al.
Published: (2025)
by: Guo, Mingning, et al.
Published: (2025)
Perception-based Image Denoising via Generative Compression
by: Nguyen, Nam, et al.
Published: (2026)
by: Nguyen, Nam, et al.
Published: (2026)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
by: Chen, Pengtao, et al.
Published: (2025)
by: Chen, Pengtao, et al.
Published: (2025)
AgriChat: A Multimodal Large Language Model for Agriculture Image Understanding
by: Boudiaf, Abderrahmene, et al.
Published: (2026)
by: Boudiaf, Abderrahmene, et al.
Published: (2026)
Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning
by: Sun, Penglei, et al.
Published: (2025)
by: Sun, Penglei, et al.
Published: (2025)
Enhancing Shape Perception and Segmentation Consistency for Industrial Image Inspection
by: Mao, Guoxuan, et al.
Published: (2025)
by: Mao, Guoxuan, et al.
Published: (2025)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
by: Hyeon-Woo, Nam, et al.
Published: (2024)
by: Hyeon-Woo, Nam, et al.
Published: (2024)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
by: Cho, Jang Hyun, et al.
Published: (2025)
by: Cho, Jang Hyun, et al.
Published: (2025)
AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning
by: Xu, Binxiao, et al.
Published: (2026)
by: Xu, Binxiao, et al.
Published: (2026)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
by: Guo, Zhongbin, et al.
Published: (2025)
by: Guo, Zhongbin, et al.
Published: (2025)
Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
by: Xu, Yifan, et al.
Published: (2025)
by: Xu, Yifan, et al.
Published: (2025)
Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2025)
by: Li, Yuan, et al.
Published: (2025)
Decoupling Perception and Calibration: Label-Efficient Image Quality Assessment Framework
by: Li, Xinyue, et al.
Published: (2026)
by: Li, Xinyue, et al.
Published: (2026)
MediSee: Reasoning-based Pixel-level Perception in Medical Images
by: Tong, Qinyue, et al.
Published: (2025)
by: Tong, Qinyue, et al.
Published: (2025)
Similar Items
-
Explaining Low Perception Model Competency with High-Competency Counterfactuals
by: Pohland, Sara, et al.
Published: (2025) -
Competency-Aware Planning for Probabilistically Safe Navigation Under Perception Uncertainty
by: Pohland, Sara, et al.
Published: (2024) -
PaRCE: Probabilistic and Reconstruction-based Competency Estimation for CNN-based Image Classification
by: Pohland, Sara, et al.
Published: (2024) -
A Deep Learning Approach for Augmenting Perceptional Understanding of Histopathology Images
by: Hu, Xiaoqian
Published: (2025) -
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
by: Fang, Tianchen, et al.
Published: (2025)