Visual Enumeration Remains Challenging for Multimodal Generative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Testolin, Alberto, Hou, Kuinan, Zorzi, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Evolutionary Network Architecture Search Framework with Adaptive Multimodal Fusion for Hand Gesture Recognition
by: Xia, Yizhang, et al.
Published: (2024)
by: Xia, Yizhang, et al.
Published: (2024)
Estimating the distribution of numerosity and non-numerical visual magnitudes in natural scenes using computer vision
by: Hou, Kuinan, et al.
Published: (2024)
by: Hou, Kuinan, et al.
Published: (2024)
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
by: Hou, Kuinan, et al.
Published: (2025)
by: Hou, Kuinan, et al.
Published: (2025)
NeuralDiffuser: Neuroscience-inspired Diffusion Guidance for fMRI Visual Reconstruction
by: Li, Haoyu, et al.
Published: (2024)
by: Li, Haoyu, et al.
Published: (2024)
EPRBench: A High-Quality Benchmark Dataset for Event Stream Based Visual Place Recognition
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
Long-Term Visual Object Tracking with Event Cameras: An Associative Memory Augmented Tracker and A Benchmark Dataset
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
SoftReMish: A Novel Activation Function for Enhanced Convolutional Neural Networks for Visual Recognition Performance
by: Gücen, Mustafa Bayram
Published: (2025)
by: Gücen, Mustafa Bayram
Published: (2025)
Elastic Spiking Transformers for Efficient Gesture Understanding
by: Ancilotto, Alberto, et al.
Published: (2026)
by: Ancilotto, Alberto, et al.
Published: (2026)
EvoCAD: Evolutionary CAD Code Generation with Vision Language Models
by: Preintner, Tobias, et al.
Published: (2025)
by: Preintner, Tobias, et al.
Published: (2025)
Structure Abstraction and Generalization in a Hippocampal-Entorhinal Inspired World Model
by: Zhang, Tianqiu, et al.
Published: (2026)
by: Zhang, Tianqiu, et al.
Published: (2026)
Motion Illusions Generated Using Predictive Neural Networks Also Fool Humans
by: Sinapayen, Lana, et al.
Published: (2021)
by: Sinapayen, Lana, et al.
Published: (2021)
Mice to Machines: Neural Representations from Visual Cortex for Domain Generalization
by: Qazi, Ahmed, et al.
Published: (2025)
by: Qazi, Ahmed, et al.
Published: (2025)
Toward Generalized Detection of Synthetic Media: Limitations, Challenges, and the Path to Multimodal Solutions
by: Hussain, Redwan, et al.
Published: (2025)
by: Hussain, Redwan, et al.
Published: (2025)
Investigating the generative dynamics of energy-based neural networks
by: Tausani, Lorenzo, et al.
Published: (2023)
by: Tausani, Lorenzo, et al.
Published: (2023)
A developmental approach for training deep belief networks
by: Zambra, Matteo, et al.
Published: (2022)
by: Zambra, Matteo, et al.
Published: (2022)
Event Stream based Human Action Recognition: A High-Definition Benchmark Dataset and Algorithms
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
CADE: Cosine Annealing Differential Evolution for Spiking Neural Network
by: Jiang, Runhua, et al.
Published: (2024)
by: Jiang, Runhua, et al.
Published: (2024)
Enhanced Temporal Processing in Spiking Neural Networks for Static Object Detection Using 3D Convolutions
by: He, Huaxu
Published: (2024)
by: He, Huaxu
Published: (2024)
VELoRA: A Low-Rank Adaptation Approach for Efficient RGB-Event based Recognition
by: Chen, Lan, et al.
Published: (2024)
by: Chen, Lan, et al.
Published: (2024)
CCSRP: Robust Pruning of Spiking Neural Networks through Cooperative Coevolution
by: Song, Zichen, et al.
Published: (2024)
by: Song, Zichen, et al.
Published: (2024)
SNN-PAR: Energy Efficient Pedestrian Attribute Recognition via Spiking Neural Networks
by: Wang, Haiyang, et al.
Published: (2024)
by: Wang, Haiyang, et al.
Published: (2024)
Neural Echos: Depthwise Convolutional Filters Replicate Biological Receptive Fields
by: Babaiee, Zahra, et al.
Published: (2024)
by: Babaiee, Zahra, et al.
Published: (2024)
VFM-Det: Towards High-Performance Vehicle Detection via Large Foundation Models
by: Wu, Wentao, et al.
Published: (2024)
by: Wu, Wentao, et al.
Published: (2024)
LM-HT SNN: Enhancing the Performance of SNN to ANN Counterpart through Learnable Multi-hierarchical Threshold Model
by: Hao, Zecheng, et al.
Published: (2024)
by: Hao, Zecheng, et al.
Published: (2024)
EAS-SNN: End-to-End Adaptive Sampling and Representation for Event-based Detection with Recurrent Spiking Neural Networks
by: Wang, Ziming, et al.
Published: (2024)
by: Wang, Ziming, et al.
Published: (2024)
Faster and Stronger: When ANN-SNN Conversion Meets Parallel Spiking Calculation
by: Hao, Zecheng, et al.
Published: (2024)
by: Hao, Zecheng, et al.
Published: (2024)
Retain, Blend, and Exchange: A Quality-aware Spatial-Stereo Fusion Approach for Event Stream Recognition
by: Chen, Lan, et al.
Published: (2024)
by: Chen, Lan, et al.
Published: (2024)
VerifIoU -- Robustness of Object Detection to Perturbations
by: Cohen, Noémie, et al.
Published: (2024)
by: Cohen, Noémie, et al.
Published: (2024)
Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
by: Lei, Zhenxin, et al.
Published: (2024)
by: Lei, Zhenxin, et al.
Published: (2024)
Self-supervised cross-modality learning for uncertainty-aware object detection and recognition in applications which lack pre-labelled training data
by: Mehboob, Irum, et al.
Published: (2024)
by: Mehboob, Irum, et al.
Published: (2024)
Impacts of Darwinian Evolution on Pre-trained Deep Neural Networks
by: Du, Guodong, et al.
Published: (2024)
by: Du, Guodong, et al.
Published: (2024)
Going beyond Compositions, DDPMs Can Produce Zero-Shot Interpolations
by: Deschenaux, Justin, et al.
Published: (2024)
by: Deschenaux, Justin, et al.
Published: (2024)
QKFormer: Hierarchical Spiking Transformer using Q-K Attention
by: Zhou, Chenlin, et al.
Published: (2024)
by: Zhou, Chenlin, et al.
Published: (2024)
Evaluating Driver Readiness in Conditionally Automated Vehicles from Eye-Tracking Data and Head Pose
by: Kazemi, Mostafa, et al.
Published: (2024)
by: Kazemi, Mostafa, et al.
Published: (2024)
A Tournament of Transformation Models: B-Spline-based vs. Mesh-based Multi-Objective Deformable Image Registration
by: Andreadis, Georgios, et al.
Published: (2024)
by: Andreadis, Georgios, et al.
Published: (2024)
SpikingRTNH: Spiking Neural Network for 4D Radar Object Detection
by: Paek, Dong-Hee, et al.
Published: (2025)
by: Paek, Dong-Hee, et al.
Published: (2025)
From Neurons to Computation: Biological Reservoir Computing for Pattern Recognition
by: Iannello, Ludovico, et al.
Published: (2025)
by: Iannello, Ludovico, et al.
Published: (2025)
From Neural Activity to Computation: Biological Reservoirs for Pattern Recognition in Digit Classification
by: Iannello, Ludovico, et al.
Published: (2025)
by: Iannello, Ludovico, et al.
Published: (2025)
Patch of Invisibility: Naturalistic Physical Black-Box Adversarial Attacks on Object Detectors
by: Lapid, Raz, et al.
Published: (2023)
by: Lapid, Raz, et al.
Published: (2023)
Revisiting Color-Event based Tracking: A Unified Network, Dataset, and Metric
by: Tang, Chuanming, et al.
Published: (2022)
by: Tang, Chuanming, et al.
Published: (2022)
Similar Items
-
An Evolutionary Network Architecture Search Framework with Adaptive Multimodal Fusion for Hand Gesture Recognition
by: Xia, Yizhang, et al.
Published: (2024) -
Estimating the distribution of numerosity and non-numerical visual magnitudes in natural scenes using computer vision
by: Hou, Kuinan, et al.
Published: (2024) -
Assessing the Visual Enumeration Abilities of Specialized Counting Architectures and Vision-Language Models
by: Hou, Kuinan, et al.
Published: (2025) -
NeuralDiffuser: Neuroscience-inspired Diffusion Guidance for fMRI Visual Reconstruction
by: Li, Haoyu, et al.
Published: (2024) -
EPRBench: A High-Quality Benchmark Dataset for Event Stream Based Visual Place Recognition
by: Wang, Xiao, et al.
Published: (2026)