Text-promptable Object Counting via Quantity Awareness Enhancement
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Miaojing, Zhang, Xiaowen, Yue, Zijie, Luo, Yong, Zhao, Cairong, Li, Li |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
by: Zhang, Xiaowen, et al.
Published: (2026)
by: Zhang, Xiaowen, et al.
Published: (2026)
Enhancing Space-time Video Super-resolution via Spatial-temporal Feature Interaction
by: Yue, Zijie, et al.
Published: (2022)
by: Yue, Zijie, et al.
Published: (2022)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026)
by: Shi, Miaojing, et al.
Published: (2026)
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
by: Yuan, Linfeng, et al.
Published: (2023)
by: Yuan, Linfeng, et al.
Published: (2023)
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
by: Xue, Yuhao, et al.
Published: (2025)
by: Xue, Yuhao, et al.
Published: (2025)
FAAR: Efficient Frequency-Aware Multi-Task Fine-Tuning via Automatic Rank Selection
by: Fontana, Maxime, et al.
Published: (2026)
by: Fontana, Maxime, et al.
Published: (2026)
Boosting Object Detection with Zero-Shot Day-Night Domain Adaptation
by: Du, Zhipeng, et al.
Published: (2023)
by: Du, Zhipeng, et al.
Published: (2023)
Large Model driven Radiology Report Generation with Clinical Quality Reinforcement Learning
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
Bootstrapping Vision-language Models for Self-supervised Remote Physiological Measurement
by: Yue, Zijie, et al.
Published: (2024)
by: Yue, Zijie, et al.
Published: (2024)
CountCluster: Training-Free Object Quantity Guidance with Cross-Attention Map Clustering for Text-to-Image Generation
by: Lee, Joohyeon, et al.
Published: (2025)
by: Lee, Joohyeon, et al.
Published: (2025)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
by: Zhang, Da, et al.
Published: (2026)
by: Zhang, Da, et al.
Published: (2026)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
Text Promptable Surgical Instrument Segmentation with Vision-Language Models
by: Zhou, Zijian, et al.
Published: (2023)
by: Zhou, Zijian, et al.
Published: (2023)
UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
Quality and Quantity: Unveiling a Million High-Quality Images for Text-to-Image Synthesis in Fashion Design
by: Yu, Jia, et al.
Published: (2023)
by: Yu, Jia, et al.
Published: (2023)
SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
AdaTreeFormer: Few Shot Domain Adaptation for Tree Counting from a Single High-Resolution Image
by: Amirkolaee, Hamed Amini, et al.
Published: (2024)
by: Amirkolaee, Hamed Amini, et al.
Published: (2024)
Instruction Guided Multi Object Image Editing with Quantity and Layout Consistency
by: Tan, Jiaqi, et al.
Published: (2025)
by: Tan, Jiaqi, et al.
Published: (2025)
CPDM: Content-Preserving Diffusion Model for Underwater Image Enhancement
by: Shi, Xiaowen, et al.
Published: (2024)
by: Shi, Xiaowen, et al.
Published: (2024)
Memory-guided Network with Uncertainty-based Feature Augmentation for Few-shot Semantic Segmentation
by: Chen, Xinyue, et al.
Published: (2024)
by: Chen, Xinyue, et al.
Published: (2024)
OpenPSG: Open-set Panoptic Scene Graph Generation via Large Multimodal Models
by: Zhou, Zijian, et al.
Published: (2024)
by: Zhou, Zijian, et al.
Published: (2024)
LineCounter: Learning Handwritten Text Line Segmentation by Counting
by: Li, Deng, et al.
Published: (2021)
by: Li, Deng, et al.
Published: (2021)
Multitask Learning in Minimally Invasive Surgical Vision: A Review
by: Alabi, Oluwatosin, et al.
Published: (2024)
by: Alabi, Oluwatosin, et al.
Published: (2024)
VLPrompt: Vision-Language Prompting for Panoptic Scene Graph Generation
by: Zhou, Zijian, et al.
Published: (2023)
by: Zhou, Zijian, et al.
Published: (2023)
Grounding Surgical Action Triplets with Instrument Instance Segmentation: A Dataset and Target-Aware Fusion Approach
by: Alabi, Oluwatosin, et al.
Published: (2025)
by: Alabi, Oluwatosin, et al.
Published: (2025)
MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping
by: Gao, Junyao, et al.
Published: (2026)
by: Gao, Junyao, et al.
Published: (2026)
Self-Creative Text-to-Object Generation using Semantic-Aware Spatial Weighting
by: Yu, Yue, et al.
Published: (2026)
by: Yu, Yue, et al.
Published: (2026)
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP
by: Zhang, Zeliang, et al.
Published: (2024)
by: Zhang, Zeliang, et al.
Published: (2024)
One Object, Multiple Lies: A Benchmark for Cross-task Adversarial Attack on Unified Vision-Language Models
by: Zhao, Jiale, et al.
Published: (2025)
by: Zhao, Jiale, et al.
Published: (2025)
Online Video Quality Enhancement with Spatial-Temporal Look-up Tables
by: Qu, Zefan, et al.
Published: (2023)
by: Qu, Zefan, et al.
Published: (2023)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
An unsupervised approach towards promptable defect segmentation in laser-based additive manufacturing by Segment Anything
by: Era, Israt Zarin, et al.
Published: (2023)
by: Era, Israt Zarin, et al.
Published: (2023)
Beyond Quantity: Distribution-Aware Labeling for Visual Grounding
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Towards the Influence of Text Quantity on Writer Retrieval
by: Peer, Marco, et al.
Published: (2025)
by: Peer, Marco, et al.
Published: (2025)
LIME-Eval: Rethinking Low-light Image Enhancement Evaluation via Object Detection
by: Li, Mingjia, et al.
Published: (2024)
by: Li, Mingjia, et al.
Published: (2024)
Counting Stacked Objects
by: Dumery, Corentin, et al.
Published: (2024)
by: Dumery, Corentin, et al.
Published: (2024)
Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge Transfer
by: Chen, Xinyue, et al.
Published: (2024)
by: Chen, Xinyue, et al.
Published: (2024)
DiffPhysBA: Diffusion-based Physical Backdoor Attack against Person Re-Identification in Real-World
by: Sun, Wenli, et al.
Published: (2024)
by: Sun, Wenli, et al.
Published: (2024)
CamPVG: Camera-Controlled Panoramic Video Generation with Epipolar-Aware Diffusion
by: Ji, Chenhao, et al.
Published: (2025)
by: Ji, Chenhao, et al.
Published: (2025)
Similar Items
-
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
by: Zhang, Xiaowen, et al.
Published: (2026) -
Enhancing Space-time Video Super-resolution via Spatial-temporal Feature Interaction
by: Yue, Zijie, et al.
Published: (2022) -
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
by: Shi, Miaojing, et al.
Published: (2026) -
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
by: Yuan, Linfeng, et al.
Published: (2023) -
TRAIL: Transferable Robust Adversarial Images via Latent diffusion
by: Xue, Yuhao, et al.
Published: (2025)