Unified Dense Prediction of Video Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Lehan, Qi, Lu, Li, Xiangtai, Li, Sheng, Jampani, Varun, Yang, Ming-Hsuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Video Prediction Transformers without Recurrence or Convolution
von: Tang, Yujin, et al.
Veröffentlicht: (2024)
von: Tang, Yujin, et al.
Veröffentlicht: (2024)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
von: Hu, Tao, et al.
Veröffentlicht: (2026)
von: Hu, Tao, et al.
Veröffentlicht: (2026)
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024)
von: Qi, Lu, et al.
Veröffentlicht: (2024)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
von: Kilian, Maciej, et al.
Veröffentlicht: (2024)
DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)
VRMDiff: Text-Guided Video Referring Matting Generation of Diffusion
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
von: Yang, Lehan, et al.
Veröffentlicht: (2025)
Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Exploiting Diffusion Prior for Generalizable Dense Prediction
von: Lee, Hsin-Ying, et al.
Veröffentlicht: (2023)
von: Lee, Hsin-Ying, et al.
Veröffentlicht: (2023)
Reason3D: Searching and Reasoning 3D Segmentation via Large Language Model
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
von: Huang, Kuan-Chih, et al.
Veröffentlicht: (2024)
Reference Twice: A Simple and Unified Baseline for Few-Shot Instance Segmentation
von: Han, Yue, et al.
Veröffentlicht: (2023)
von: Han, Yue, et al.
Veröffentlicht: (2023)
Human Video Generation from a Single Image with 3D Pose and View Control
von: Wang, Tiantian, et al.
Veröffentlicht: (2026)
von: Wang, Tiantian, et al.
Veröffentlicht: (2026)
Dense360: Dense Understanding from Omnidirectional Panoramas
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
von: Zhou, Yikang, et al.
Veröffentlicht: (2025)
SemFlow: Binding Semantic Segmentation and Image Synthesis via Rectified Flow
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
von: Xu, Shilin, et al.
Veröffentlicht: (2024)
Stable Part Diffusion 4D: Multi-View RGB and Kinematic Parts Video Generation
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Hao, et al.
Veröffentlicht: (2025)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
von: Zhang, Junyi, et al.
Veröffentlicht: (2023)
von: Zhang, Junyi, et al.
Veröffentlicht: (2023)
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
von: Bao, Jingzhi, et al.
Veröffentlicht: (2024)
von: Bao, Jingzhi, et al.
Veröffentlicht: (2024)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
PanopticPartFormer++: A Unified and Decoupled View for Panoptic Part Segmentation
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
von: Li, Xiangtai, et al.
Veröffentlicht: (2023)
SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D Generation
von: Yao, Chun-Han, et al.
Veröffentlicht: (2025)
von: Yao, Chun-Han, et al.
Veröffentlicht: (2025)
Pyramid Diffusion for Fine 3D Large Scene Generation
von: Liu, Yuheng, et al.
Veröffentlicht: (2023)
von: Liu, Yuheng, et al.
Veröffentlicht: (2023)
Effective Adapter for Face Recognition in the Wild
von: Liu, Yunhao, et al.
Veröffentlicht: (2023)
von: Liu, Yunhao, et al.
Veröffentlicht: (2023)
Mamba or RWKV: Exploring High-Quality and High-Efficiency Segment Anything Model
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
von: Yuan, Haobo, et al.
Veröffentlicht: (2024)
Decouple and Track: Benchmarking and Improving Video Diffusion Transformers for Motion Transfer
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
von: Shi, Qingyu, et al.
Veröffentlicht: (2025)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
von: Zhang, Junyi, et al.
Veröffentlicht: (2024)
von: Zhang, Junyi, et al.
Veröffentlicht: (2024)
Human-in-Context: Unified Cross-Domain 3D Human Motion Modeling via In-Context Learning
von: Liu, Mengyuan, et al.
Veröffentlicht: (2025)
von: Liu, Mengyuan, et al.
Veröffentlicht: (2025)
RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
von: Yue, Jingtong, et al.
Veröffentlicht: (2025)
von: Yue, Jingtong, et al.
Veröffentlicht: (2025)
4th PVUW MeViS 3rd Place Report: Sa2VA
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
Tuning-Free Image Editing with Fidelity and Editability via Unified Latent Diffusion Model
von: Mao, Qi, et al.
Veröffentlicht: (2025)
von: Mao, Qi, et al.
Veröffentlicht: (2025)
ConDense: Consistent 2D/3D Pre-training for Dense and Sparse Features from Multi-View Images
von: Zhang, Xiaoshuai, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaoshuai, et al.
Veröffentlicht: (2024)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
von: Nguyen, Hieu T., et al.
Veröffentlicht: (2024)
von: Nguyen, Hieu T., et al.
Veröffentlicht: (2024)
A Simple Approach to Unifying Diffusion-based Conditional Generation
von: Li, Xirui, et al.
Veröffentlicht: (2024)
von: Li, Xirui, et al.
Veröffentlicht: (2024)
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
von: He, Jixuan, et al.
Veröffentlicht: (2025)
von: He, Jixuan, et al.
Veröffentlicht: (2025)
SMooDi: Stylized Motion Diffusion Model
von: Zhong, Lei, et al.
Veröffentlicht: (2024)
von: Zhong, Lei, et al.
Veröffentlicht: (2024)
CoCo4D: Comprehensive and Complex 4D Scene Generation
von: Zhou, Junwei, et al.
Veröffentlicht: (2025)
von: Zhou, Junwei, et al.
Veröffentlicht: (2025)
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
von: Lin, Yuanze, et al.
Veröffentlicht: (2025)
Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint
von: Zhou, Junwei, et al.
Veröffentlicht: (2024)
von: Zhou, Junwei, et al.
Veröffentlicht: (2024)
SyncNoise: Geometrically Consistent Noise Prediction for Text-based 3D Scene Editing
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
von: Li, Ruihuang, et al.
Veröffentlicht: (2024)
Frequency Domain-Based Diffusion Model for Unpaired Image Dehazing
von: Liu, Chengxu, et al.
Veröffentlicht: (2025)
von: Liu, Chengxu, et al.
Veröffentlicht: (2025)
Learning Deblurring Texture Prior from Unpaired Data with Diffusion Model
von: Liu, Chengxu, et al.
Veröffentlicht: (2025)
von: Liu, Chengxu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Video Prediction Transformers without Recurrence or Convolution
von: Tang, Yujin, et al.
Veröffentlicht: (2024) -
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
von: Hu, Tao, et al.
Veröffentlicht: (2026) -
Generalizable Entity Grounding via Assistance of Large Language Model
von: Qi, Lu, et al.
Veröffentlicht: (2024) -
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
von: Kilian, Maciej, et al.
Veröffentlicht: (2024) -
DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency
von: Qi, Mengshi, et al.
Veröffentlicht: (2025)