Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
Fuente:
arXiv
Saved in:
| Main Authors: | Xuan, Shiyu, Li, Zechao, Tang, Jinhui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
by: Xuan, Shiyu, et al.
Published: (2026)
by: Xuan, Shiyu, et al.
Published: (2026)
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024)
by: Tian, Changyao, et al.
Published: (2024)
Adaptive Perception for Unified Visual Multi-modal Object Tracking
by: Hu, Xiantao, et al.
Published: (2025)
by: Hu, Xiantao, et al.
Published: (2025)
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
by: Zhang, Zijian, et al.
Published: (2024)
by: Zhang, Zijian, et al.
Published: (2024)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Unified Multi-modal Diagnostic Framework with Reconstruction Pre-training and Heterogeneity-combat Tuning
by: Zhang, Yupei, et al.
Published: (2024)
by: Zhang, Yupei, et al.
Published: (2024)
Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
by: Zhu, Zixin, et al.
Published: (2024)
by: Zhu, Zixin, et al.
Published: (2024)
See the Text: From Tokenization to Visual Reading
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
DVF: Advancing Robust and Accurate Fine-Grained Image Retrieval with Retrieval Guidelines
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
by: Dong, Zhe, et al.
Published: (2025)
by: Dong, Zhe, et al.
Published: (2025)
DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Multi-modal Vision Pre-training for Medical Image Analysis
by: Rui, Shaohao, et al.
Published: (2024)
by: Rui, Shaohao, et al.
Published: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
by: Ma, Shuailei, et al.
Published: (2023)
by: Ma, Shuailei, et al.
Published: (2023)
Awesome Multi-modal Object Tracking
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
by: Liu, Haowei, et al.
Published: (2024)
by: Liu, Haowei, et al.
Published: (2024)
A Recover-then-Discriminate Framework for Robust Anomaly Detection
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
MuDPT: Multi-modal Deep-symphysis Prompt Tuning for Large Pre-trained Vision-Language Models
by: Miao, Yongzhu, et al.
Published: (2023)
by: Miao, Yongzhu, et al.
Published: (2023)
InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
by: Luo, Zihao, et al.
Published: (2025)
by: Luo, Zihao, et al.
Published: (2025)
Vision Model Pre-training on Interleaved Image-Text Data via Latent Compression Learning
by: Yang, Chenyu, et al.
Published: (2024)
by: Yang, Chenyu, et al.
Published: (2024)
GeoMM: On Geodesic Perspective for Multi-modal Learning
by: Mei, Shibin, et al.
Published: (2025)
by: Mei, Shibin, et al.
Published: (2025)
UniQA: Unified Vision-Language Pre-training for Image Quality and Aesthetic Assessment
by: Zhou, Hantao, et al.
Published: (2024)
by: Zhou, Hantao, et al.
Published: (2024)
Multi-modal Multi-task Pre-training for Improved Point Cloud Understanding
by: Liu, Liwen, et al.
Published: (2025)
by: Liu, Liwen, et al.
Published: (2025)
ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models
by: Höllein, Lukas, et al.
Published: (2024)
by: Höllein, Lukas, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
by: Wang, Jiaqi, et al.
Published: (2025)
by: Wang, Jiaqi, et al.
Published: (2025)
LayerDiff: Exploring Text-guided Multi-layered Composable Image Synthesis via Layer-Collaborative Diffusion Model
by: Huang, Runhui, et al.
Published: (2024)
by: Huang, Runhui, et al.
Published: (2024)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
by: Wang, Jiepeng, et al.
Published: (2025)
by: Wang, Jiepeng, et al.
Published: (2025)
MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
by: Liang, Qian, et al.
Published: (2025)
by: Liang, Qian, et al.
Published: (2025)
Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
by: Li, Xuchen, et al.
Published: (2024)
by: Li, Xuchen, et al.
Published: (2024)
Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
by: Jia, Shukun, et al.
Published: (2024)
by: Jia, Shukun, et al.
Published: (2024)
MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
by: Wei, Zhichao, et al.
Published: (2024)
by: Wei, Zhichao, et al.
Published: (2024)
FTMoMamba: Motion Generation with Frequency and Text State Space Models
by: Li, Chengjian, et al.
Published: (2024)
by: Li, Chengjian, et al.
Published: (2024)
PGP-DiffSR: Phase-Guided Progressive Pruning for Efficient Diffusion-based Image Super-Resolution
by: Yang, Zhongbao, et al.
Published: (2025)
by: Yang, Zhongbao, et al.
Published: (2025)
Adapting Multi-modal Large Language Model to Concept Drift From Pre-training Onwards
by: Yang, Xiaoyu, et al.
Published: (2024)
by: Yang, Xiaoyu, et al.
Published: (2024)
Diff-Tracker: Text-to-Image Diffusion Models are Unsupervised Trackers
by: Zhang, Zhengbo, et al.
Published: (2024)
by: Zhang, Zhengbo, et al.
Published: (2024)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
by: Wu, Junfeng, et al.
Published: (2024)
by: Wu, Junfeng, et al.
Published: (2024)
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
by: Zhang, Zhengbo, et al.
Published: (2026)
by: Zhang, Zhengbo, et al.
Published: (2026)
CSGO: Content-Style Composition in Text-to-Image Generation
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
by: Chen, Lan, et al.
Published: (2025)
by: Chen, Lan, et al.
Published: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
Similar Items
-
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
by: Xuan, Shiyu, et al.
Published: (2026) -
MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer
by: Tian, Changyao, et al.
Published: (2024) -
Adaptive Perception for Unified Visual Multi-modal Object Tracking
by: Hu, Xiantao, et al.
Published: (2025) -
Integrating Text and Image Pre-training for Multi-modal Algorithmic Reasoning
by: Zhang, Zijian, et al.
Published: (2024) -
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
by: Tang, Hao, et al.
Published: (2024)