Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhengbo, Tu, Zhigang, Yuan, Junsong, Soh, De Wen, Du, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Informative Sample Selection Model for Skeleton-based Action Recognition with Limited Training Samples
by: Tu, Zhigang, et al.
Published: (2025)
by: Tu, Zhigang, et al.
Published: (2025)
FADE: A Dataset for Detecting Falling Objects around Buildings in Video
by: Tu, Zhigang, et al.
Published: (2024)
by: Tu, Zhigang, et al.
Published: (2024)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
Visual Prompting for One-shot Controllable Video Editing without Inversion
by: Zhang, Zhengbo, et al.
Published: (2025)
by: Zhang, Zhengbo, et al.
Published: (2025)
UAV-OVO: Out-of-Viewpoint Generalization in UAV Action Recognition
by: Xia, Yu, et al.
Published: (2026)
by: Xia, Yu, et al.
Published: (2026)
Frequency-Controlled Diffusion Model for Versatile Text-Guided Image-to-Image Translation
by: Gao, Xiang, et al.
Published: (2024)
by: Gao, Xiang, et al.
Published: (2024)
Frequency-Enhanced Diffusion Models: Curriculum-Guided Semantic Alignment for Zero-Shot Skeleton Action Recognition
by: Zhou, Yuxi, et al.
Published: (2026)
by: Zhou, Yuxi, et al.
Published: (2026)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction
by: Zhang, Mengfei, et al.
Published: (2026)
by: Zhang, Mengfei, et al.
Published: (2026)
Unsupervised Modality Adaptation with Text-to-Image Diffusion Models for Semantic Segmentation
by: Xia, Ruihao, et al.
Published: (2024)
by: Xia, Ruihao, et al.
Published: (2024)
Aligning Instance-Semantic Sparse Representation towards Unsupervised Object Segmentation and Shape Abstraction with Repeatable Primitives
by: Li, Jiaxin, et al.
Published: (2025)
by: Li, Jiaxin, et al.
Published: (2025)
Fully Unsupervised Self-debiasing of Text-to-Image Diffusion Models
by: Vardhana, Korada Sri, et al.
Published: (2025)
by: Vardhana, Korada Sri, et al.
Published: (2025)
VSC: Visual Search Compositional Text-to-Image Diffusion Model
by: Dat, Do Huu, et al.
Published: (2025)
by: Dat, Do Huu, et al.
Published: (2025)
GeoRemover: Removing Objects and Their Causal Visual Artifacts
by: Zhu, Zixin, et al.
Published: (2025)
by: Zhu, Zixin, et al.
Published: (2025)
DiffusionTrack: Diffusion Model For Multi-Object Tracking
by: Luo, Run, et al.
Published: (2023)
by: Luo, Run, et al.
Published: (2023)
MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
by: Xiao, Changcheng, et al.
Published: (2024)
by: Xiao, Changcheng, et al.
Published: (2024)
Towards Effective and Efficient Adversarial Defense with Diffusion Models for Robust Visual Tracking
by: Xu, Long, et al.
Published: (2025)
by: Xu, Long, et al.
Published: (2025)
Customizing Text-to-Image Diffusion with Object Viewpoint Control
by: Kumari, Nupur, et al.
Published: (2024)
by: Kumari, Nupur, et al.
Published: (2024)
MotionTrack: Learning Motion Predictor for Multiple Object Tracking
by: Xiao, Changcheng, et al.
Published: (2023)
by: Xiao, Changcheng, et al.
Published: (2023)
MAEDiff: Masked Autoencoder-enhanced Diffusion Models for Unsupervised Anomaly Detection in Brain Images
by: Xu, Rui, et al.
Published: (2024)
by: Xu, Rui, et al.
Published: (2024)
VastTrack: Vast Category Visual Object Tracking
by: Peng, Liang, et al.
Published: (2024)
by: Peng, Liang, et al.
Published: (2024)
Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
by: Li, Sifan, et al.
Published: (2025)
by: Li, Sifan, et al.
Published: (2025)
Explicit Visual Prompts for Visual Object Tracking
by: Shi, Liangtao, et al.
Published: (2024)
by: Shi, Liangtao, et al.
Published: (2024)
Visual Concept-driven Image Generation with Text-to-Image Diffusion Model
by: Rahman, Tanzila, et al.
Published: (2024)
by: Rahman, Tanzila, et al.
Published: (2024)
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
by: Yin, Wen, et al.
Published: (2025)
by: Yin, Wen, et al.
Published: (2025)
Object-Conditioned Energy-Based Attention Map Alignment in Text-to-Image Diffusion Models
by: Zhang, Yasi, et al.
Published: (2024)
by: Zhang, Yasi, et al.
Published: (2024)
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
by: Shuai, Xincheng, et al.
Published: (2024)
by: Shuai, Xincheng, et al.
Published: (2024)
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
by: Zeng, Guanning, et al.
Published: (2025)
by: Zeng, Guanning, et al.
Published: (2025)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Investigating Fine- and Coarse-grained Structural Correspondences Between Deep Neural Networks and Human Object Image Similarity Judgments Using Unsupervised Alignment
by: Takahashi, Soh, et al.
Published: (2025)
by: Takahashi, Soh, et al.
Published: (2025)
QASA: Quality-Guided K-Adaptive Slot Attention for Unsupervised Object-Centric Learning
by: Ouyang, Tianran, et al.
Published: (2026)
by: Ouyang, Tianran, et al.
Published: (2026)
BEVUDA++: Geometric-aware Unsupervised Domain Adaptation for Multi-View 3D Object Detection
by: Zhang, Rongyu, et al.
Published: (2025)
by: Zhang, Rongyu, et al.
Published: (2025)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis
by: Wang, Yuanrui, et al.
Published: (2025)
by: Wang, Yuanrui, et al.
Published: (2025)
Contrastive Prompts Improve Disentanglement in Text-to-Image Diffusion Models
by: Wu, Chen, et al.
Published: (2024)
by: Wu, Chen, et al.
Published: (2024)
VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
by: Yuan, Jiayi, et al.
Published: (2026)
by: Yuan, Jiayi, et al.
Published: (2026)
HEAP: Unsupervised Object Discovery and Localization with Contrastive Grouping
by: Zhang, Xin, et al.
Published: (2023)
by: Zhang, Xin, et al.
Published: (2023)
DeTrack: In-model Latent Denoising Learning for Visual Object Tracking
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Pluralistic Salient Object Detection
by: Feng, Xuelu, et al.
Published: (2024)
by: Feng, Xuelu, et al.
Published: (2024)
Similar Items
-
Informative Sample Selection Model for Skeleton-based Action Recognition with Limited Training Samples
by: Tu, Zhigang, et al.
Published: (2025) -
FADE: A Dataset for Detecting Falling Objects around Buildings in Video
by: Tu, Zhigang, et al.
Published: (2024) -
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026) -
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024) -
Visual Prompting for One-shot Controllable Video Editing without Inversion
by: Zhang, Zhengbo, et al.
Published: (2025)