QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Wenfang, Du, Yingjun, Liu, Gaowen, Zheng, Yefeng, Snoek, Cees G. M. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Training-Free Semantic Segmentation via LLM-Supervision
by: Sun, Wenfang, et al.
Published: (2024)
by: Sun, Wenfang, et al.
Published: (2024)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026)
by: Sun, Wenfang, et al.
Published: (2026)
Union-over-Intersections: Object Detection beyond Winner-Takes-All
by: Bhowmik, Aritra, et al.
Published: (2023)
by: Bhowmik, Aritra, et al.
Published: (2023)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
by: Du, Yingjun, et al.
Published: (2023)
by: Du, Yingjun, et al.
Published: (2023)
An Image is Worth More Than 16x16 Patches: Exploring Transformers on Individual Pixels
by: Nguyen, Duy-Kien, et al.
Published: (2024)
by: Nguyen, Duy-Kien, et al.
Published: (2024)
SelEx: Self-Expertise in Fine-Grained Generalized Category Discovery
by: Rastegar, Sarah, et al.
Published: (2024)
by: Rastegar, Sarah, et al.
Published: (2024)
Segment Any 3D-Part in a Scene from a Sentence
by: Wu, Hongyu, et al.
Published: (2025)
by: Wu, Hongyu, et al.
Published: (2025)
Geometric Neural Process Fields
by: Yin, Wenzhe, et al.
Published: (2025)
by: Yin, Wenzhe, et al.
Published: (2025)
Unaligning Everything: Or Aligning Any Text to Any Image in Multimodal Models
by: Salman, Shaeke, et al.
Published: (2024)
by: Salman, Shaeke, et al.
Published: (2024)
Any-Shift Prompting for Generalization over Distributions
by: Xiao, Zehao, et al.
Published: (2024)
by: Xiao, Zehao, et al.
Published: (2024)
Continual Hyperbolic Learning of Instances and Classes
by: Ayoughi, Melika, et al.
Published: (2025)
by: Ayoughi, Melika, et al.
Published: (2025)
Spatial Reasoners for Continuous Variables in Any Domain
by: Pogodzinski, Bart, et al.
Published: (2025)
by: Pogodzinski, Bart, et al.
Published: (2025)
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
by: Nguyen, Duy-Kien, et al.
Published: (2023)
by: Nguyen, Duy-Kien, et al.
Published: (2023)
Low-Resource Vision Challenges for Foundation Models
by: Zhang, Yunhua, et al.
Published: (2024)
by: Zhang, Yunhua, et al.
Published: (2024)
Forget Vectors at Play: Universal Input Perturbations Driving Machine Unlearning in Image Classification
by: Sun, Changchang, et al.
Published: (2024)
by: Sun, Changchang, et al.
Published: (2024)
PIN: Positional Insert Unlocks Object Localisation Abilities in VLMs
by: Dorkenwald, Michael, et al.
Published: (2024)
by: Dorkenwald, Michael, et al.
Published: (2024)
Improving Long-Text Alignment for Text-to-Image Diffusion Models
by: Liu, Luping, et al.
Published: (2024)
by: Liu, Luping, et al.
Published: (2024)
Self-Evaluation Unlocks Any-Step Text-to-Image Generation
by: Yu, Xin, et al.
Published: (2025)
by: Yu, Xin, et al.
Published: (2025)
Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs
by: Narang, Arya
Published: (2025)
by: Narang, Arya
Published: (2025)
Fill the Gap: Quantifying and Reducing the Modality Gap in Image-Text Representation Learning
by: Role, François, et al.
Published: (2025)
by: Role, François, et al.
Published: (2025)
Redefining Normal: A Novel Object-Level Approach for Multi-Object Novelty Detection
by: Salehi, Mohammadreza, et al.
Published: (2024)
by: Salehi, Mohammadreza, et al.
Published: (2024)
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
Variational Bayesian Last Layers
by: Harrison, James, et al.
Published: (2024)
by: Harrison, James, et al.
Published: (2024)
Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation
by: Franchi, Gianni, et al.
Published: (2024)
by: Franchi, Gianni, et al.
Published: (2024)
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
by: Zhang, Wenlun, et al.
Published: (2026)
by: Zhang, Wenlun, et al.
Published: (2026)
Structurally Prune Anything: Any Architecture, Any Framework, Any Time
by: Wang, Xun, et al.
Published: (2024)
by: Wang, Xun, et al.
Published: (2024)
Instance-Warp: Saliency Guided Image Warping for Unsupervised Domain Adaptation
by: Zheng, Shen, et al.
Published: (2024)
by: Zheng, Shen, et al.
Published: (2024)
A Novel Convolution and Attention Mechanism-based Model for 6D Object Pose Estimation
by: Du, Alexander, et al.
Published: (2024)
by: Du, Alexander, et al.
Published: (2024)
Beyond Coarse-Grained Matching in Video-Text Retrieval
by: Chen, Aozhu, et al.
Published: (2024)
by: Chen, Aozhu, et al.
Published: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
by: Liu, Huabin, et al.
Published: (2025)
by: Liu, Huabin, et al.
Published: (2025)
Progressive Compositionality in Text-to-Image Generative Models
by: Han, Evans Xu, et al.
Published: (2024)
by: Han, Evans Xu, et al.
Published: (2024)
Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
by: Xue, Shuchen, et al.
Published: (2025)
by: Xue, Shuchen, et al.
Published: (2025)
Cocktail: Mixing Multi-Modality Controls for Text-Conditional Image Generation
by: Hu, Minghui, et al.
Published: (2023)
by: Hu, Minghui, et al.
Published: (2023)
Domain-Specialized Object Detection via Model-Level Mixtures of Experts
by: Pavlitska, Svetlana, et al.
Published: (2026)
by: Pavlitska, Svetlana, et al.
Published: (2026)
Understanding Matrix Function Normalizations in Covariance Pooling through the Lens of Riemannian Geometry
by: Chen, Ziheng, et al.
Published: (2024)
by: Chen, Ziheng, et al.
Published: (2024)
TAO-Amodal: A Benchmark for Tracking Any Object Amodally
by: Hsieh, Cheng-Yen, et al.
Published: (2023)
by: Hsieh, Cheng-Yen, et al.
Published: (2023)
Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector
by: Fu, Yuqian, et al.
Published: (2024)
by: Fu, Yuqian, et al.
Published: (2024)
Similar Items
-
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024) -
Prompt Diffusion Robustifies Any-Modality Prompt Learning
by: Du, Yingjun, et al.
Published: (2024) -
Training-Free Semantic Segmentation via LLM-Supervision
by: Sun, Wenfang, et al.
Published: (2024) -
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
by: Sun, Wenfang, et al.
Published: (2026) -
Union-over-Intersections: Object Detection beyond Winner-Takes-All
by: Bhowmik, Aritra, et al.
Published: (2023)