CountFormer: A Transformer Framework for Learning Visual Repetition and Structure in Class-Agnostic Object Counting
Fuente:
arXiv
Saved in:
| Main Authors: | Hossain, Md Tanvir, Islam, Akif, Ameen, Mohd Ruhul |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting AI-Generated Images via Diffusion Snap-Back Reconstruction: A Forensic Approach
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
by: Ameen, Mohd Ruhul, et al.
Published: (2025)
CountFormer: Multi-View Crowd Counting Transformer
by: Mo, Hong, et al.
Published: (2024)
by: Mo, Hong, et al.
Published: (2024)
Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
by: Ope, Momen Khandoker, et al.
Published: (2025)
by: Ope, Momen Khandoker, et al.
Published: (2025)
Automated Wicket-Taking Delivery Segmentation and Trajectory-Based Dismissal-Zone Analysis in Cricket Videos Using OCR-Guided YOLOv8
by: Karmoker, Joy, et al.
Published: (2025)
by: Karmoker, Joy, et al.
Published: (2025)
From Pixels to People: Satellite-Based Mapping and Quantification of Riverbank Erosion and Lost Villages in Bangladesh
by: Rafat, M Saifuzzaman, et al.
Published: (2025)
by: Rafat, M Saifuzzaman, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning for Low-Resource Languages: A Comparative Study of LLMs for Bengali Hate Speech Detection
by: Islam, Akif, et al.
Published: (2025)
by: Islam, Akif, et al.
Published: (2025)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
GroundCount: Grounding Vision-Language Models with Object Detection for Mitigating Counting Hallucinations
by: Chen, Boyuan, et al.
Published: (2026)
by: Chen, Boyuan, et al.
Published: (2026)
Mutually-Aware Feature Learning for Few-Shot Object Counting
by: Jeon, Yerim, et al.
Published: (2024)
by: Jeon, Yerim, et al.
Published: (2024)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models
by: Alzahrani, Reem, et al.
Published: (2026)
by: Alzahrani, Reem, et al.
Published: (2026)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
IVAC-P2L: Leveraging Irregular Repetition Priors for Improving Video Action Counting
by: Wang, Hang, et al.
Published: (2024)
by: Wang, Hang, et al.
Published: (2024)
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
CountQA: How Well Do MLLMs Count in the Wild?
by: Tamarapalli, Jayant Sravan, et al.
Published: (2025)
by: Tamarapalli, Jayant Sravan, et al.
Published: (2025)
RCCFormer: A Robust Crowd Counting Network Based on Transformer
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
MultiCounter: Multiple Action Agnostic Repetition Counting in Untrimmed Videos
by: Tang, Yin, et al.
Published: (2024)
by: Tang, Yin, et al.
Published: (2024)
Make It Count: Text-to-Image Generation with an Accurate Number of Objects
by: Binyamin, Lital, et al.
Published: (2024)
by: Binyamin, Lital, et al.
Published: (2024)
Seed Kernel Counting using Domain Randomization and Object Tracking Neural Networks
by: Margapuri, Venkat, et al.
Published: (2023)
by: Margapuri, Venkat, et al.
Published: (2023)
TGBFormer: Transformer-GraphFormer Blender Network for Video Object Detection
by: Qi, Qiang, et al.
Published: (2025)
by: Qi, Qiang, et al.
Published: (2025)
OCCAM: Class-Agnostic, Training-Free, Prior-Free and Multi-Class Object Counting
by: Spanakis, Michail, et al.
Published: (2026)
by: Spanakis, Michail, et al.
Published: (2026)
Counting Through Occlusion: Framework for Open World Amodal Counting
by: Arib, Safaeid Hossain, et al.
Published: (2025)
by: Arib, Safaeid Hossain, et al.
Published: (2025)
Mamba-MOC: A Multicategory Remote Object Counting via State Space Model
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
by: Che, Liwei, et al.
Published: (2026)
by: Che, Liwei, et al.
Published: (2026)
Curriculum for Crowd Counting -- Is it Worthy?
by: Khan, Muhammad Asif, et al.
Published: (2024)
by: Khan, Muhammad Asif, et al.
Published: (2024)
Count What You Want: Exemplar Identification and Few-shot Counting of Human Actions in the Wild
by: Huang, Yifeng, et al.
Published: (2023)
by: Huang, Yifeng, et al.
Published: (2023)
Versatile Incremental Learning: Towards Class and Domain-Agnostic Incremental Learning
by: Park, Min-Yeong, et al.
Published: (2024)
by: Park, Min-Yeong, et al.
Published: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
by: Guo, Xuyang, et al.
Published: (2025)
by: Guo, Xuyang, et al.
Published: (2025)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
Balancing Interpretability and Performance in Motor Imagery EEG Classification: A Comparative Study of ANFIS-FBCSP-PSO and EEGNet
by: Aktar, Farjana, et al.
Published: (2025)
by: Aktar, Farjana, et al.
Published: (2025)
Temporally-Aware Supervised Contrastive Learning for Polyp Counting in Colonoscopy
by: Parolari, Luca, et al.
Published: (2025)
by: Parolari, Luca, et al.
Published: (2025)
SITUATE -- Synthetic Object Counting Dataset for VLM training
by: Peinl, René, et al.
Published: (2026)
by: Peinl, René, et al.
Published: (2026)
Personalized Federated Segmentation with Shared Feature Aggregation and Boundary-Focused Calibration
by: Tashdeed, Ishmam, et al.
Published: (2025)
by: Tashdeed, Ishmam, et al.
Published: (2025)
Does it Really Count? Assessing Semantic Grounding in Text-Guided Class-Agnostic Counting
by: Pacini, Giacomo, et al.
Published: (2026)
by: Pacini, Giacomo, et al.
Published: (2026)
CountingFruit: Language-Guided 3D Fruit Counting with Semantic Gaussian Splatting
by: Li, Fengze, et al.
Published: (2025)
by: Li, Fengze, et al.
Published: (2025)
DEEGITS: Deep Learning based Framework for Measuring Heterogenous Traffic State in Challenging Traffic Scenarios
by: Islam, Muttahirul, et al.
Published: (2024)
by: Islam, Muttahirul, et al.
Published: (2024)
Bound Tightening Network for Robust Crowd Counting
by: Wu, Qiming
Published: (2024)
by: Wu, Qiming
Published: (2024)
Similar Items
-
Detecting AI-Generated Images via Diffusion Snap-Back Reconstruction: A Forensic Approach
by: Ameen, Mohd Ruhul, et al.
Published: (2025) -
Towards a Humanized Social-Media Ecosystem: AI-Augmented HCI Design Patterns for Safety, Agency & Well-Being
by: Ameen, Mohd Ruhul, et al.
Published: (2025) -
CountFormer: Multi-View Crowd Counting Transformer
by: Mo, Hong, et al.
Published: (2024) -
Oitijjo-3D: Generative AI Framework for Rapid 3D Heritage Reconstruction from Street View Imagery
by: Ope, Momen Khandoker, et al.
Published: (2025) -
Automated Wicket-Taking Delivery Segmentation and Trajectory-Based Dismissal-Zone Analysis in Cricket Videos Using OCR-Guided YOLOv8
by: Karmoker, Joy, et al.
Published: (2025)