CountGD: Multi-Modal Open-World Counting
Fuente:
arXiv
Saved in:
| Main Authors: | Amini-Naieni, Niki, Han, Tengda, Zisserman, Andrew |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025)
by: Amini-Naieni, Niki, et al.
Published: (2025)
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
by: Dumery, Corentin, et al.
Published: (2026)
by: Dumery, Corentin, et al.
Published: (2026)
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024)
by: Keser, Mert, et al.
Published: (2024)
Learning to Count without Annotations
by: Knobel, Lukas, et al.
Published: (2023)
by: Knobel, Lukas, et al.
Published: (2023)
Instant Uncertainty Calibration of NeRFs Using a Meta-Calibrator
by: Amini-Naieni, Niki, et al.
Published: (2023)
by: Amini-Naieni, Niki, et al.
Published: (2023)
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
by: Perrett, Toby, et al.
Published: (2024)
by: Perrett, Toby, et al.
Published: (2024)
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
by: Xie, Junyu, et al.
Published: (2026)
by: Xie, Junyu, et al.
Published: (2026)
OVR: A Dataset for Open Vocabulary Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
Counting Through Occlusion: Framework for Open World Amodal Counting
by: Arib, Safaeid Hossain, et al.
Published: (2025)
by: Arib, Safaeid Hossain, et al.
Published: (2025)
Character-Centric Understanding of Animated Movies
by: Gui, Zhongrui, et al.
Published: (2025)
by: Gui, Zhongrui, et al.
Published: (2025)
Seeing without Pixels: Perception from Camera Trajectories
by: Xue, Zihui, et al.
Published: (2025)
by: Xue, Zihui, et al.
Published: (2025)
Benchmarking Vision Foundation Models for Input Monitoring in Autonomous Driving
by: Keser, Mert, et al.
Published: (2025)
by: Keser, Mert, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
Multi-modal Crowd Counting via Modal Emulation
by: Wang, Chenhao, et al.
Published: (2024)
by: Wang, Chenhao, et al.
Published: (2024)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
Multi-modal Crowd Counting via a Broker Modality
by: Meng, Haoliang, et al.
Published: (2024)
by: Meng, Haoliang, et al.
Published: (2024)
A Short Note on Evaluating RepNet for Temporal Repetition Counting in Videos
by: Dwibedi, Debidatta, et al.
Published: (2024)
by: Dwibedi, Debidatta, et al.
Published: (2024)
CountFormer: Multi-View Crowd Counting Transformer
by: Mo, Hong, et al.
Published: (2024)
by: Mo, Hong, et al.
Published: (2024)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
by: Zhao, Yiming, et al.
Published: (2025)
by: Zhao, Yiming, et al.
Published: (2025)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
FocalCount: Towards Class-Count Imbalance in Class-Agnostic Counting
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
CountDiffusion: Text-to-Image Synthesis with Training-Free Counting-Guidance Diffusion
by: Li, Yanyu, et al.
Published: (2025)
by: Li, Yanyu, et al.
Published: (2025)
CountMamba: Exploring Multi-directional Selective State-Space Models for Plant Counting
by: He, Hulingxiao, et al.
Published: (2024)
by: He, Hulingxiao, et al.
Published: (2024)
Learning from Streaming Video with Orthogonal Gradients
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond
by: Wu, Guanyao, et al.
Published: (2025)
by: Wu, Guanyao, et al.
Published: (2025)
The Manga Whisperer: Automatically Generating Transcriptions for Comics
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
by: Korbar, Bruno, et al.
Published: (2025)
by: Korbar, Bruno, et al.
Published: (2025)
Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
by: Bagad, Piyush, et al.
Published: (2025)
by: Bagad, Piyush, et al.
Published: (2025)
From Panels to Prose: Generating Literary Narratives from Comics
by: Sachdeva, Ragav, et al.
Published: (2025)
by: Sachdeva, Ragav, et al.
Published: (2025)
Joint Counting, Detection and Re-Identification for Multi-Object Tracking
by: Ren, Weihong, et al.
Published: (2022)
by: Ren, Weihong, et al.
Published: (2022)
Rethinking Cell Counting Methods: Decoupling Counting and Localization
by: Zheng, Zixuan, et al.
Published: (2025)
by: Zheng, Zixuan, et al.
Published: (2025)
RS-OVC: Open-Vocabulary Counting for Remote-Sensing Data
by: Shor, Tamir, et al.
Published: (2026)
by: Shor, Tamir, et al.
Published: (2026)
Unified Open-World Segmentation with Multi-Modal Prompts
by: Liu, Yang, et al.
Published: (2025)
by: Liu, Yang, et al.
Published: (2025)
Multi-Modal Prototypes for Open-World Semantic Segmentation
by: Yang, Yuhuan, et al.
Published: (2023)
by: Yang, Yuhuan, et al.
Published: (2023)
A Survey on Class-Agnostic Counting: Advancements from Reference-Based to Open-World Text-Guided Approaches
by: Ciampi, Luca, et al.
Published: (2025)
by: Ciampi, Luca, et al.
Published: (2025)
Count Anything
by: Lei, Mengqi, et al.
Published: (2026)
by: Lei, Mengqi, et al.
Published: (2026)
Decoupling What to Count and Where to See for Referring Expression Counting
by: Zou, Yuda, et al.
Published: (2025)
by: Zou, Yuda, et al.
Published: (2025)
Similar Items
-
CountGD++: Generalized Prompting for Open-World Counting
by: Amini-Naieni, Niki, et al.
Published: (2025) -
Open-World Object Counting in Videos
by: Amini-Naieni, Niki, et al.
Published: (2025) -
The MixCount Dataset: Bridging the Data Gap for Open-Vocabulary Object Counting
by: Dumery, Corentin, et al.
Published: (2026) -
Unveiling Ontological Commitment in Multi-Modal Foundation Models
by: Keser, Mert, et al.
Published: (2024) -
Learning to Count without Annotations
by: Knobel, Lukas, et al.
Published: (2023)