VLCounter: Text-aware Visual Representation for Zero-Shot Object Counting
Fuente:
arXiv
Saved in:
| Main Authors: | Kang, Seunggu, Moon, WonJun, Kim, Euiyeon, Heo, Jae-Pil |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026)
by: Seong, Hyun Seok, et al.
Published: (2026)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
by: Moon, WonJun, et al.
Published: (2023)
by: Moon, WonJun, et al.
Published: (2023)
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
by: Jeon, Yerim, et al.
Published: (2025)
by: Jeon, Yerim, et al.
Published: (2025)
Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Progressive Proxy Anchor Propagation for Unsupervised Semantic Segmentation
by: Seong, Hyun Seok, et al.
Published: (2024)
by: Seong, Hyun Seok, et al.
Published: (2024)
Mitigating Background Shift in Class-Incremental Semantic Segmentation
by: Park, Gilhan, et al.
Published: (2024)
by: Park, Gilhan, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
Mutually-Aware Feature Learning for Few-Shot Object Counting
by: Jeon, Yerim, et al.
Published: (2024)
by: Jeon, Yerim, et al.
Published: (2024)
Looking Beyond the Window: Global-Local Aligned CLIP for Training-free Open-Vocabulary Semantic Segmentation
by: Lee, ByeongCheol, et al.
Published: (2026)
by: Lee, ByeongCheol, et al.
Published: (2026)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
GSGAN: Adversarial Learning for Hierarchical Generation of 3D Gaussian Splats
by: Hyun, Sangeek, et al.
Published: (2024)
by: Hyun, Sangeek, et al.
Published: (2024)
Diversity-aware Channel Pruning for StyleGAN Compression
by: Chung, Jiwoo, et al.
Published: (2024)
by: Chung, Jiwoo, et al.
Published: (2024)
Activating Self-Attention for Multi-Scene Absolute Pose Regression
by: Lee, Miso, et al.
Published: (2024)
by: Lee, Miso, et al.
Published: (2024)
Long-term Pre-training for Temporal Action Detection with Transformers
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models
by: Koh, Eunseo, et al.
Published: (2025)
by: Koh, Eunseo, et al.
Published: (2025)
Noise-free Optimization in Early Training Steps for Image Super-Resolution
by: Lee, MinKyu, et al.
Published: (2023)
by: Lee, MinKyu, et al.
Published: (2023)
Disambiguating 2D-3D Correspondences in Gaussian Splatting-based Feature Fields for Visual Localization
by: Lee, Miso, et al.
Published: (2026)
by: Lee, Miso, et al.
Published: (2026)
Boundary-Recovering Network for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Style Injection in Diffusion: A Training-free Approach for Adapting Large-scale Diffusion Models for Style Transfer
by: Chung, Jiwoo, et al.
Published: (2023)
by: Chung, Jiwoo, et al.
Published: (2023)
Expanding Zero-Shot Object Counting with Rich Prompts
by: Zhu, Huilin, et al.
Published: (2025)
by: Zhu, Huilin, et al.
Published: (2025)
AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
by: Jin, Woojeong, et al.
Published: (2026)
by: Jin, Woojeong, et al.
Published: (2026)
Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot Segmentation
by: Park, Suho, et al.
Published: (2025)
by: Park, Suho, et al.
Published: (2025)
Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation
by: Chung, Jiwoo, et al.
Published: (2025)
by: Chung, Jiwoo, et al.
Published: (2025)
Temporally Consistent Long-Term Memory for 3D Single Object Tracking
by: Yoo, Jaejoon, et al.
Published: (2026)
by: Yoo, Jaejoon, et al.
Published: (2026)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
by: Zhang, Da, et al.
Published: (2026)
by: Zhang, Da, et al.
Published: (2026)
Text-guided Zero-Shot Object Localization
by: Wang, Jingjing, et al.
Published: (2024)
by: Wang, Jingjing, et al.
Published: (2024)
Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
by: Heo, Jun-Woo, et al.
Published: (2026)
by: Heo, Jun-Woo, et al.
Published: (2026)
Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization
by: Lee, MinKyu, et al.
Published: (2025)
by: Lee, MinKyu, et al.
Published: (2025)
Analogical Trajectory Transfer
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
Prediction-Feedback DETR for Temporal Action Detection
by: Kim, Jihwan, et al.
Published: (2024)
by: Kim, Jihwan, et al.
Published: (2024)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
by: Cho, CH, et al.
Published: (2025)
by: Cho, CH, et al.
Published: (2025)
White Aggregation and Restoration for Few-shot 3D Point Cloud Semantic Segmentation
by: Im, Jiyun, et al.
Published: (2025)
by: Im, Jiyun, et al.
Published: (2025)
Cross-scale Aligned Supervision for Training GANs
by: Hyun, Sangeek, et al.
Published: (2026)
by: Hyun, Sangeek, et al.
Published: (2026)
CountZES: Counting via Zero-Shot Exemplar Selection
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
by: Siddiqui, Muhammad Ibraheem, et al.
Published: (2025)
Zero-Shot Aerial Object Detection with Visual Description Regularization
by: Zang, Zhengqing, et al.
Published: (2024)
by: Zang, Zhengqing, et al.
Published: (2024)
Zero-Shot Object-Centric Representation Learning
by: Didolkar, Aniket, et al.
Published: (2024)
by: Didolkar, Aniket, et al.
Published: (2024)
Similar Items
-
From Vicious to Virtuous Cycles: Synergistic Representation Learning for Unsupervised Video Object-Centric Learning
by: Seong, Hyun Seok, et al.
Published: (2026) -
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026) -
Selective Contrastive Learning for Weakly Supervised Affordance Grounding
by: Moon, WonJun, et al.
Published: (2025) -
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding
by: Moon, WonJun, et al.
Published: (2023) -
Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
by: Jeon, Yerim, et al.
Published: (2025)