Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Gia Khanh, Huang, Yifeng, Hoai, Minh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CountEx: Fine-Grained Counting via Exemplars and Exclusion
by: Huang, Yifeng, et al.
Published: (2026)
by: Huang, Yifeng, et al.
Published: (2026)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023)
by: Huang, Yifeng, et al.
Published: (2023)
Decoupling What to Count and Where to See for Referring Expression Counting
by: Zou, Yuda, et al.
Published: (2025)
by: Zou, Yuda, et al.
Published: (2025)
From Sora What We Can See: A Survey of Text-to-Video Generation
by: Sun, Rui, et al.
Published: (2024)
by: Sun, Rui, et al.
Published: (2024)
DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
by: Huang, Yifeng, et al.
Published: (2025)
by: Huang, Yifeng, et al.
Published: (2025)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
by: Upadhyay, Ujjwal, et al.
Published: (2025)
by: Upadhyay, Ujjwal, et al.
Published: (2025)
Aligning What EEG Can See: Structural Representations for Brain-Vision Matching
by: Tang, Jingyi, et al.
Published: (2026)
by: Tang, Jingyi, et al.
Published: (2026)
What You See is What You Ask: Evaluating Audio Descriptions
by: Kala, Divy, et al.
Published: (2025)
by: Kala, Divy, et al.
Published: (2025)
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Detecting Precise Hand Touch Moments in Egocentric Video
by: Nguyen, Huy Anh, et al.
Published: (2026)
by: Nguyen, Huy Anh, et al.
Published: (2026)
HOIST-Former: Hand-held Objects Identification, Segmentation, and Tracking in the Wild
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
by: Narasimhaswamy, Supreeth, et al.
Published: (2024)
What Are We Measuring When We Evaluate Large Vision-Language Models? An Analysis of Latent Factors and Biases
by: Tiong, Anthony Meng Huat, et al.
Published: (2024)
by: Tiong, Anthony Meng Huat, et al.
Published: (2024)
FrameOracle: Learning What to See and How Much to See in Videos
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
What Does It Mean for a Medical AI System to Be Right?
by: Gitau, Antony
Published: (2026)
by: Gitau, Antony
Published: (2026)
Dual Strategies for Test-Time Adaptation
by: Phuong, Nam Nguyen, et al.
Published: (2026)
by: Phuong, Nam Nguyen, et al.
Published: (2026)
Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented Augmentation
by: Corvi, Riccardo, et al.
Published: (2025)
by: Corvi, Riccardo, et al.
Published: (2025)
MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
by: Nguyen, Duc Duy, et al.
Published: (2026)
by: Nguyen, Duc Duy, et al.
Published: (2026)
Detecting Omissions in Geographic Maps through Computer Vision
by: Nguyen, Phuc D. A., et al.
Published: (2024)
by: Nguyen, Phuc D. A., et al.
Published: (2024)
What Can We Learn from Harry Potter? An Exploratory Study of Visual Representation Learning from Atypical Videos
by: Sun, Qiyue, et al.
Published: (2025)
by: Sun, Qiyue, et al.
Published: (2025)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
Pixel Motion Diffusion is What We Need for Robot Control
by: Nguyen, E-Ro, et al.
Published: (2025)
by: Nguyen, E-Ro, et al.
Published: (2025)
Driver Attention Tracking and Analysis
by: Nguyen, Dat Viet Thanh, et al.
Published: (2024)
by: Nguyen, Dat Viet Thanh, et al.
Published: (2024)
Where Am I and What Will I See: An Auto-Regressive Model for Spatial Localization and View Prediction
by: Chen, Junyi, et al.
Published: (2024)
by: Chen, Junyi, et al.
Published: (2024)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
by: Liu, Chengxin, et al.
Published: (2026)
by: Liu, Chengxin, et al.
Published: (2026)
Underwater Image Enhancement with Physical-based Denoising Diffusion Implicit Models
by: Bach, Nguyen Gia, et al.
Published: (2024)
by: Bach, Nguyen Gia, et al.
Published: (2024)
Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models
by: Ho, Nhan, et al.
Published: (2026)
by: Ho, Nhan, et al.
Published: (2026)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
by: Pei, Gensheng, et al.
Published: (2025)
by: Pei, Gensheng, et al.
Published: (2025)
Seeing What Shouldn't Be There: Counterfactual GANs for Medical Image Attribution
by: Murtaza, Shakeeb
Published: (2026)
by: Murtaza, Shakeeb
Published: (2026)
Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
What You See is (Usually) What You Get: Multimodal Prototype Networks that Abstain from Expensive Modalities
by: Bahng, Muchang, et al.
Published: (2025)
by: Bahng, Muchang, et al.
Published: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
by: Song, Junha, et al.
Published: (2026)
by: Song, Junha, et al.
Published: (2026)
Scale What Counts, Mask What Matters: Evaluating Foundation Models for Zero-Shot Cross-Domain Wi-Fi Sensing
by: Jiang, Cheng, et al.
Published: (2025)
by: Jiang, Cheng, et al.
Published: (2025)
What Color Is It? A Text-Interference Multimodal Hallucination Benchmark
by: Zhao, Jinkun, et al.
Published: (2025)
by: Zhao, Jinkun, et al.
Published: (2025)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
PairAug: What Can Augmented Image-Text Pairs Do for Radiology?
by: Xie, Yutong, et al.
Published: (2024)
by: Xie, Yutong, et al.
Published: (2024)
See What You Seek: Semantic Contextual Integration for Cloth-Changing Person Re-Identification
by: Han, Xiyu, et al.
Published: (2024)
by: Han, Xiyu, et al.
Published: (2024)
xAI-CV: An Overview of Explainable Artificial Intelligence in Computer Vision
by: Van Tu, Nguyen, et al.
Published: (2025)
by: Van Tu, Nguyen, et al.
Published: (2025)
Can You Trust What You See? Alpha Channel No-Box Attacks on Video Object Detection
by: Yi, Ariana, et al.
Published: (2025)
by: Yi, Ariana, et al.
Published: (2025)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
by: Poletti, Silvia, et al.
Published: (2026)
by: Poletti, Silvia, et al.
Published: (2026)
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
by: Sengupta, Saurav, et al.
Published: (2025)
by: Sengupta, Saurav, et al.
Published: (2025)
Similar Items
-
CountEx: Fine-Grained Counting via Exemplars and Exclusion
by: Huang, Yifeng, et al.
Published: (2026) -
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023) -
Decoupling What to Count and Where to See for Referring Expression Counting
by: Zou, Yuda, et al.
Published: (2025) -
From Sora What We Can See: A Survey of Text-to-Video Generation
by: Sun, Rui, et al.
Published: (2024) -
DualMat: PBR Material Estimation via Coherent Dual-Path Diffusion
by: Huang, Yifeng, et al.
Published: (2025)