Exploring Pre-trained Text-to-Video Diffusion Models for Referring Video Object Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Zixin, Feng, Xuelu, Chen, Dongdong, Yuan, Junsong, Qiao, Chunming, Hua, Gang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pluralistic Salient Object Detection
von: Feng, Xuelu, et al.
Veröffentlicht: (2024)
von: Feng, Xuelu, et al.
Veröffentlicht: (2024)
GeoRemover: Removing Objects and Their Causal Visual Artifacts
von: Zhu, Zixin, et al.
Veröffentlicht: (2025)
von: Zhu, Zixin, et al.
Veröffentlicht: (2025)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
Benchmarking Large and Small MLLMs
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)
Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
von: Luan, Tianyu, et al.
Veröffentlicht: (2025)
von: Luan, Tianyu, et al.
Veröffentlicht: (2025)
SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
von: Liu, Sheng, et al.
Veröffentlicht: (2025)
Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
von: Chen, Jingxi, et al.
Veröffentlicht: (2024)
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
von: Yuan, Linfeng, et al.
Veröffentlicht: (2023)
Weakly-Supervised Referring Video Object Segmentation through Text Supervision
von: Shi, Miaojing, et al.
Veröffentlicht: (2026)
von: Shi, Miaojing, et al.
Veröffentlicht: (2026)
EventRR: Event Referential Reasoning for Referring Video Object Segmentation
von: Xu, Huihui, et al.
Veröffentlicht: (2025)
von: Xu, Huihui, et al.
Veröffentlicht: (2025)
Leveraging Text-to-Image Diffusion Models for Unsupervised Visual Object Tracking
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
von: Wang, Hengkang, et al.
Veröffentlicht: (2025)
DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyeonwoo, et al.
Veröffentlicht: (2025)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
von: Choi, Sun-Hyuk, et al.
Veröffentlicht: (2025)
von: Choi, Sun-Hyuk, et al.
Veröffentlicht: (2025)
EndoMamba: An Efficient Foundation Model for Endoscopic Videos via Hierarchical Pre-training
von: Tian, Qingyao, et al.
Veröffentlicht: (2025)
von: Tian, Qingyao, et al.
Veröffentlicht: (2025)
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
von: Gao, Mingqi, et al.
Veröffentlicht: (2026)
DiffRIS: Enhancing Referring Remote Sensing Image Segmentation with Pre-trained Text-to-Image Diffusion Models
von: Dong, Zhe, et al.
Veröffentlicht: (2025)
von: Dong, Zhe, et al.
Veröffentlicht: (2025)
InterRVOS: Interaction-aware Referring Video Object Segmentation
von: Jin, Woojeong, et al.
Veröffentlicht: (2025)
von: Jin, Woojeong, et al.
Veröffentlicht: (2025)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory
von: Miao, Bo, et al.
Veröffentlicht: (2024)
von: Miao, Bo, et al.
Veröffentlicht: (2024)
Referring Video Object Segmentation with Cross-Modality Proxy Queries
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
von: Sun, Baoli, et al.
Veröffentlicht: (2025)
Extreme Video Compression with Pre-trained Diffusion Models
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation
von: Yarram, Sudhir, et al.
Veröffentlicht: (2024)
von: Yarram, Sudhir, et al.
Veröffentlicht: (2024)
FADE: A Dataset for Detecting Falling Objects around Buildings in Video
von: Tu, Zhigang, et al.
Veröffentlicht: (2024)
von: Tu, Zhigang, et al.
Veröffentlicht: (2024)
Continual Forgetting for Pre-trained Vision Models
von: Zhao, Hongbo, et al.
Veröffentlicht: (2024)
von: Zhao, Hongbo, et al.
Veröffentlicht: (2024)
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
von: Lee, Seunghun, et al.
Veröffentlicht: (2025)
SVAC: Scaling Is All You Need For Referring Video Object Segmentation
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Referring Video Object Segmentation via Language-aligned Track Selection
von: Kim, Seongchan, et al.
Veröffentlicht: (2024)
von: Kim, Seongchan, et al.
Veröffentlicht: (2024)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
Zero-Shot Video Semantic Segmentation based on Pre-Trained Diffusion Models
von: Wang, Qian, et al.
Veröffentlicht: (2024)
von: Wang, Qian, et al.
Veröffentlicht: (2024)
OmnimatteZero: Fast Training-free Omnimatte with Pre-trained Video Diffusion Models
von: Samuel, Dvir, et al.
Veröffentlicht: (2025)
von: Samuel, Dvir, et al.
Veröffentlicht: (2025)
RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation
von: Li, Yonglin, et al.
Veröffentlicht: (2023)
von: Li, Yonglin, et al.
Veröffentlicht: (2023)
Enhancing Sa2VA for Referent Video Object Segmentation: 2nd Solution for 7th LSVOS RVOS Track
von: Hong, Ran, et al.
Veröffentlicht: (2025)
von: Hong, Ran, et al.
Veröffentlicht: (2025)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
von: Li, Changzhen, et al.
Veröffentlicht: (2025)
von: Li, Changzhen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Pluralistic Salient Object Detection
von: Feng, Xuelu, et al.
Veröffentlicht: (2024) -
GeoRemover: Removing Objects and Their Causal Visual Artifacts
von: Zhu, Zixin, et al.
Veröffentlicht: (2025) -
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
von: Feng, Xuelu, et al.
Veröffentlicht: (2025) -
Benchmarking Large and Small MLLMs
von: Feng, Xuelu, et al.
Veröffentlicht: (2025) -
Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model
von: Zhang, Ruixin, et al.
Veröffentlicht: (2025)