What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
Fuente:
arXiv
Salvato in:
| Autori principali: | Xu, Guangkai, Ge, Yongtao, Liu, Mingyu, Fan, Chengxiang, Xie, Kangyang, Zhao, Zhiyue, Chen, Hao, Shen, Chunhua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Generative Video Matting
di: Ge, Yongtao, et al.
Pubblicazione: (2025)
di: Ge, Yongtao, et al.
Pubblicazione: (2025)
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
di: Zhang, Songyan, et al.
Pubblicazione: (2025)
di: Zhang, Songyan, et al.
Pubblicazione: (2025)
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
di: Ge, Yongtao, et al.
Pubblicazione: (2024)
DiffCalib: Reformulating Monocular Camera Calibration as Diffusion-Based Dense Incident Map Generation
di: He, Xiankang, et al.
Pubblicazione: (2024)
di: He, Xiankang, et al.
Pubblicazione: (2024)
Tinker: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
di: Zhao, Canyu, et al.
Pubblicazione: (2025)
Generative Active Learning for Long-tailed Instance Segmentation
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
Zippo: Zipping Color and Transparency Distributions into a Single Diffusion Model
di: Xie, Kangyang, et al.
Pubblicazione: (2024)
di: Xie, Kangyang, et al.
Pubblicazione: (2024)
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion Models
di: Wang, Wen, et al.
Pubblicazione: (2023)
di: Wang, Wen, et al.
Pubblicazione: (2023)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
di: Zhu, Muzhi, et al.
Pubblicazione: (2024)
MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data
di: Fan, Chengxiang, et al.
Pubblicazione: (2024)
di: Fan, Chengxiang, et al.
Pubblicazione: (2024)
Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
di: Xu, Guangkai, et al.
Pubblicazione: (2026)
di: Xu, Guangkai, et al.
Pubblicazione: (2026)
MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence
di: Zhao, Canyu, et al.
Pubblicazione: (2024)
di: Zhao, Canyu, et al.
Pubblicazione: (2024)
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
di: Li, Liyang, et al.
Pubblicazione: (2026)
di: Li, Liyang, et al.
Pubblicazione: (2026)
Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
di: Xu, Shaocong, et al.
Pubblicazione: (2025)
di: Xu, Shaocong, et al.
Pubblicazione: (2025)
TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking
di: Nam, Jisu, et al.
Pubblicazione: (2026)
di: Nam, Jisu, et al.
Pubblicazione: (2026)
Diff-2-in-1: Bridging Generation and Dense Perception with Diffusion Models
di: Zheng, Shuhong, et al.
Pubblicazione: (2024)
di: Zheng, Shuhong, et al.
Pubblicazione: (2024)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
di: Fan, Yue, et al.
Pubblicazione: (2024)
di: Fan, Yue, et al.
Pubblicazione: (2024)
FreeCompose: Generic Zero-Shot Image Composition with Diffusion Prior
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
di: Chen, Zhekai, et al.
Pubblicazione: (2024)
Repurposing Geometric Foundation Models for Multi-view Diffusion
di: Jang, Wooseok, et al.
Pubblicazione: (2026)
di: Jang, Wooseok, et al.
Pubblicazione: (2026)
Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation
di: Ke, Bingxin, et al.
Pubblicazione: (2023)
di: Ke, Bingxin, et al.
Pubblicazione: (2023)
Active-O3: Empowering Multimodal Large Language Models with Active Perception via GRPO
di: Zhu, Muzhi, et al.
Pubblicazione: (2025)
di: Zhu, Muzhi, et al.
Pubblicazione: (2025)
PTQAT: A Hybrid Parameter-Efficient Quantization Algorithm for 3D Perception Tasks
di: Wang, Xinhao, et al.
Pubblicazione: (2025)
di: Wang, Xinhao, et al.
Pubblicazione: (2025)
Guided Diffusion-based Generation of Adversarial Objects for Real-World Monocular Depth Estimation Attacks
di: Chen, Yongtao, et al.
Pubblicazione: (2025)
di: Chen, Yongtao, et al.
Pubblicazione: (2025)
ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
di: Wang, Kaijun, et al.
Pubblicazione: (2025)
di: Wang, Kaijun, et al.
Pubblicazione: (2025)
GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State
di: Shen, Guole, et al.
Pubblicazione: (2025)
di: Shen, Guole, et al.
Pubblicazione: (2025)
MARBLE: Multi-Aspect Reward Balance for Diffusion RL
di: Zhao, Canyu, et al.
Pubblicazione: (2026)
di: Zhao, Canyu, et al.
Pubblicazione: (2026)
Repurposing 2D Diffusion Models with Gaussian Atlas for 3D Generation
di: Xiang, Tiange, et al.
Pubblicazione: (2025)
di: Xiang, Tiange, et al.
Pubblicazione: (2025)
DiffSplat: Repurposing Image Diffusion Models for Scalable Gaussian Splat Generation
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
di: Lin, Chenguo, et al.
Pubblicazione: (2025)
Text2Stereo: Repurposing Stable Diffusion for Stereo Generation with Consistency Rewards
di: Garg, Aakash, et al.
Pubblicazione: (2025)
di: Garg, Aakash, et al.
Pubblicazione: (2025)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
di: Lin, Jiantao, et al.
Pubblicazione: (2025)
di: Lin, Jiantao, et al.
Pubblicazione: (2025)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
di: Zhong, Hao, et al.
Pubblicazione: (2025)
di: Zhong, Hao, et al.
Pubblicazione: (2025)
RFAssigner: A Generic Label Assignment Strategy for Dense Object Detection
di: Guan, Ziqian, et al.
Pubblicazione: (2026)
di: Guan, Ziqian, et al.
Pubblicazione: (2026)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
di: Huang, Junming, et al.
Pubblicazione: (2026)
di: Huang, Junming, et al.
Pubblicazione: (2026)
Task-oriented Learnable Diffusion Timesteps for Universal Few-shot Learning of Dense Tasks
di: Oh, Changgyoon, et al.
Pubblicazione: (2025)
di: Oh, Changgyoon, et al.
Pubblicazione: (2025)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation
di: Ge, Mingji, et al.
Pubblicazione: (2026)
di: Ge, Mingji, et al.
Pubblicazione: (2026)
Repurposing Video Diffusion Transformers for Robust Point Tracking
di: Son, Soowon, et al.
Pubblicazione: (2025)
di: Son, Soowon, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Generative Video Matting
di: Ge, Yongtao, et al.
Pubblicazione: (2025) -
GeoBench: Benchmarking and Analyzing Monocular Geometry Estimation Models
di: Ge, Yongtao, et al.
Pubblicazione: (2024) -
DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks
di: Zhao, Canyu, et al.
Pubblicazione: (2025) -
POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction
di: Zhang, Songyan, et al.
Pubblicazione: (2025) -
Diffusion Models are Efficient Data Generators for Human Mesh Recovery
di: Ge, Yongtao, et al.
Pubblicazione: (2024)