Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Kuinan, Mi, Jing, Zorzi, Marco, Ballan, Lamberto, Testolin, Alberto |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Enumeration Remains Challenging for Multimodal Generative AI
by: Testolin, Alberto, et al.
Published: (2024)
by: Testolin, Alberto, et al.
Published: (2024)
Sequential Enumeration in Large Language Models
by: Hou, Kuinan, et al.
Published: (2025)
by: Hou, Kuinan, et al.
Published: (2025)
Estimating the distribution of numerosity and non-numerical visual magnitudes in natural scenes using computer vision
by: Hou, Kuinan, et al.
Published: (2024)
by: Hou, Kuinan, et al.
Published: (2024)
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
by: Parolari, Luca, et al.
Published: (2024)
by: Parolari, Luca, et al.
Published: (2024)
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)
by: Pang, Xingzhou, et al.
Published: (2026)
Distilling Knowledge for Short-to-Long Term Trajectory Prediction
by: Das, Sourav, et al.
Published: (2023)
by: Das, Sourav, et al.
Published: (2023)
Towards Polyp Counting In Full-Procedure Colonoscopy Videos
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings
by: Parolari, Luca, et al.
Published: (2026)
by: Parolari, Luca, et al.
Published: (2026)
AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
by: Mai, Zheda, et al.
Published: (2025)
by: Mai, Zheda, et al.
Published: (2025)
Leveraging Vision Language Models for Specialized Agricultural Tasks
by: Arshad, Muhammad Arbab, et al.
Published: (2024)
by: Arshad, Muhammad Arbab, et al.
Published: (2024)
Multiview Progress Prediction of Robot Activities
by: Zoppellari, Elena, et al.
Published: (2026)
by: Zoppellari, Elena, et al.
Published: (2026)
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
by: Izzo, Elena, et al.
Published: (2025)
by: Izzo, Elena, et al.
Published: (2025)
Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
by: Tsaprazlis, Efthymios, et al.
Published: (2025)
You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps
by: Carraro, Riccardo, et al.
Published: (2026)
by: Carraro, Riccardo, et al.
Published: (2026)
ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models
by: Clavié, Benjamin, et al.
Published: (2025)
by: Clavié, Benjamin, et al.
Published: (2025)
Attribute-based Visual Reprogramming for Vision-Language Models
by: Cai, Chengyi, et al.
Published: (2025)
by: Cai, Chengyi, et al.
Published: (2025)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language Models
by: Berasi, Davide, et al.
Published: (2025)
by: Berasi, Davide, et al.
Published: (2025)
EasyARC: Evaluating Vision Language Models on True Visual Reasoning
by: Unsal, Mert, et al.
Published: (2025)
by: Unsal, Mert, et al.
Published: (2025)
Forecasting and Visualizing Air Quality from Sky Images with Vision-Language Models
by: Vahdatpour, Mohammad Saleh, et al.
Published: (2025)
by: Vahdatpour, Mohammad Saleh, et al.
Published: (2025)
Combining YOLO and Visual Rhythm for Vehicle Counting
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
by: Ribeiro, Victor Nascimento, et al.
Published: (2025)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
by: Zhang, Zezhou, et al.
Published: (2026)
by: Zhang, Zezhou, et al.
Published: (2026)
Automated Detection of Dolphin Whistles with Convolutional Networks and Transfer Learning
by: Korkmaz, Burla Nur, et al.
Published: (2022)
by: Korkmaz, Burla Nur, et al.
Published: (2022)
Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models
by: Li, Xu, et al.
Published: (2024)
by: Li, Xu, et al.
Published: (2024)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
by: Hartsock, Iryna, et al.
Published: (2024)
by: Hartsock, Iryna, et al.
Published: (2024)
Unleashing the Power of Vision-Language Models for Long-Tailed Multi-Label Visual Recognition
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
Visual Modality Prompt for Adapting Vision-Language Object Detectors
by: Medeiros, Heitor R., et al.
Published: (2024)
by: Medeiros, Heitor R., et al.
Published: (2024)
Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
by: Golovanevsky, Michal, et al.
Published: (2025)
by: Golovanevsky, Michal, et al.
Published: (2025)
Visual Perturbation and Adaptive Hard Negative Contrastive Learning for Compositional Reasoning in Vision-Language Models
by: Huang, Xin, et al.
Published: (2025)
by: Huang, Xin, et al.
Published: (2025)
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
by: Chu, Xu, et al.
Published: (2025)
by: Chu, Xu, et al.
Published: (2025)
TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy Exploration
by: Li, Chunxiao, et al.
Published: (2026)
by: Li, Chunxiao, et al.
Published: (2026)
Test-Time Training for Visual Foresight Vision-Language-Action Models
by: Park, Sangwu, et al.
Published: (2026)
by: Park, Sangwu, et al.
Published: (2026)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
by: Xiao, Lei, et al.
Published: (2025)
by: Xiao, Lei, et al.
Published: (2025)
Jailbreaking Vision-Language Models Through the Visual Modality
by: Azulay, Aharon, et al.
Published: (2026)
by: Azulay, Aharon, et al.
Published: (2026)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
by: Hou, Haowen, et al.
Published: (2024)
by: Hou, Haowen, et al.
Published: (2024)
Diversity Matters: Revisiting Test-Time Compute in Vision-Language Models
by: Tong, Yijie, et al.
Published: (2026)
by: Tong, Yijie, et al.
Published: (2026)
When Does Visual Prompting Outperform Linear Probing for Vision-Language Models? A Likelihood Perspective
by: Tsao, Hsi-Ai, et al.
Published: (2024)
by: Tsao, Hsi-Ai, et al.
Published: (2024)
Similar Items
-
Visual Enumeration Remains Challenging for Multimodal Generative AI
by: Testolin, Alberto, et al.
Published: (2024) -
Sequential Enumeration in Large Language Models
by: Hou, Kuinan, et al.
Published: (2025) -
Estimating the distribution of numerosity and non-numerical visual magnitudes in natural scenes using computer vision
by: Hou, Kuinan, et al.
Published: (2024) -
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
by: Parolari, Luca, et al.
Published: (2024) -
Unveiling the Visual Counting Bottleneck in Vision-Language Models
by: Pang, Xingzhou, et al.
Published: (2026)