Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | He, Jing, Li, Haodong, Yin, Wei, Liang, Yixun, Li, Leheng, Zhou, Kaiqiang, Zhang, Hongbo, Liu, Bingbing, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
by: He, Jing, et al.
Published: (2025)
by: He, Jing, et al.
Published: (2025)
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
by: Chen, Suzeyu, et al.
Published: (2026)
by: Chen, Suzeyu, et al.
Published: (2026)
Adv3D: Generating 3D Adversarial Examples for 3D Object Detection in Driving Scenarios with NeRF
by: Li, Leheng, et al.
Published: (2023)
by: Li, Leheng, et al.
Published: (2023)
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
by: Li, Leheng, et al.
Published: (2024)
by: Li, Leheng, et al.
Published: (2024)
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
by: He, Jing, et al.
Published: (2024)
by: He, Jing, et al.
Published: (2024)
BiTAA: A Bi-Task Adversarial Attack for Object Detection and Depth Estimation via 3D Gaussian Splatting
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
Exploiting Diffusion Prior for Generalizable Dense Prediction
by: Lee, Hsin-Ying, et al.
Published: (2023)
by: Lee, Hsin-Ying, et al.
Published: (2023)
Unified Dense Prediction of Video Diffusion
by: Yang, Lehan, et al.
Published: (2025)
by: Yang, Lehan, et al.
Published: (2025)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Bi-TTA: Bidirectional Test-Time Adapter for Remote Physiological Measurement
by: Li, Haodong, et al.
Published: (2024)
by: Li, Haodong, et al.
Published: (2024)
Forging Vision Foundation Models for Autonomous Driving: Challenges, Methodologies, and Opportunities
by: Yan, Xu, et al.
Published: (2024)
by: Yan, Xu, et al.
Published: (2024)
CellVTA: Enhancing Vision Foundation Models for Accurate Cell Segmentation and Classification
by: Yang, Yang, et al.
Published: (2025)
by: Yang, Yang, et al.
Published: (2025)
Frequency Dynamic Convolution for Dense Image Prediction
by: Chen, Linwei, et al.
Published: (2025)
by: Chen, Linwei, et al.
Published: (2025)
DPBridge: Latent Diffusion Bridge for Dense Prediction
by: Ji, Haorui, et al.
Published: (2024)
by: Ji, Haorui, et al.
Published: (2024)
TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation
by: Li, Mingwei, et al.
Published: (2026)
by: Li, Mingwei, et al.
Published: (2026)
BiDense: Binarization for Dense Prediction
by: Yin, Rui, et al.
Published: (2024)
by: Yin, Rui, et al.
Published: (2024)
Adversarial Patch Generation for Visual-Infrared Dense Prediction Tasks via Joint Position-Color Optimization
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
AttenST: A Training-Free Attention-Driven Style Transfer Framework with Pre-Trained Diffusion Models
by: Huang, Bo, et al.
Published: (2025)
by: Huang, Bo, et al.
Published: (2025)
FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM
by: Wu, Yuchen, et al.
Published: (2025)
by: Wu, Yuchen, et al.
Published: (2025)
Text-Audio-Visual-conditioned Diffusion Model for Video Saliency Prediction
by: Yu, Li, et al.
Published: (2025)
by: Yu, Li, et al.
Published: (2025)
ESG-Net: Event-Aware Semantic Guided Network for Dense Audio-Visual Event Localization
by: Li, Huilai, et al.
Published: (2025)
by: Li, Huilai, et al.
Published: (2025)
DreamMapping: High-Fidelity Text-to-3D Generation via Variational Distribution Mapping
by: Cai, Zeyu, et al.
Published: (2024)
by: Cai, Zeyu, et al.
Published: (2024)
Learned Image Compression with Dictionary-based Entropy Model
by: Lu, Jingbo, et al.
Published: (2025)
by: Lu, Jingbo, et al.
Published: (2025)
DOME: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model
by: Gu, Songen, et al.
Published: (2024)
by: Gu, Songen, et al.
Published: (2024)
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
Progressive Focused Transformer for Single Image Super-Resolution
by: Long, Wei, et al.
Published: (2025)
by: Long, Wei, et al.
Published: (2025)
Exploring Sparse Visual Prompt for Domain Adaptive Dense Prediction
by: Yang, Senqiao, et al.
Published: (2023)
by: Yang, Senqiao, et al.
Published: (2023)
Gaussian Shannon: High-Precision Diffusion Model Watermarking Based on Communication
by: Zhang, Yi, et al.
Published: (2026)
by: Zhang, Yi, et al.
Published: (2026)
DA$^{2}$: Depth Anything in Any Direction
by: Li, Haodong, et al.
Published: (2025)
by: Li, Haodong, et al.
Published: (2025)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
Scaling Dense Event-Stream Pretraining from Visual Foundation Models
by: Chen, Zhiwen, et al.
Published: (2026)
by: Chen, Zhiwen, et al.
Published: (2026)
DVD: Deterministic Video Depth Estimation with Generative Priors
by: Zhang, Hongfei, et al.
Published: (2026)
by: Zhang, Hongfei, et al.
Published: (2026)
3DGAA: Realistic and Robust 3D Gaussian-based Adversarial Attack for Autonomous Driving
by: Zhang, Yixun, et al.
Published: (2025)
by: Zhang, Yixun, et al.
Published: (2025)
ATD: Improved Transformer with Adaptive Token Dictionary for Image Restoration
by: Zhang, Leheng, et al.
Published: (2026)
by: Zhang, Leheng, et al.
Published: (2026)
ODTrack: Online Dense Temporal Token Learning for Visual Tracking
by: Zheng, Yaozong, et al.
Published: (2024)
by: Zheng, Yaozong, et al.
Published: (2024)
Transcending the Limit of Local Window: Advanced Super-Resolution Transformer with Adaptive Token Dictionary
by: Zhang, Leheng, et al.
Published: (2024)
by: Zhang, Leheng, et al.
Published: (2024)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Fast Inference of Visual Autoregressive Model with Adjacency-Adaptive Dynamical Draft Trees
by: Lei, Haodong, et al.
Published: (2025)
by: Lei, Haodong, et al.
Published: (2025)
Uncertainty-guided Perturbation for Image Super-Resolution Diffusion Model
by: Zhang, Leheng, et al.
Published: (2025)
by: Zhang, Leheng, et al.
Published: (2025)
Similar Items
-
Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model
by: He, Jing, et al.
Published: (2025) -
OmniBooth: Learning Latent Control for Image Synthesis with Multi-modal Instruction
by: Li, Leheng, et al.
Published: (2024) -
SPOT-Occ: Sparse Prototype-guided Transformer for Camera-based 3D Occupancy Prediction
by: Chen, Suzeyu, et al.
Published: (2026) -
Adv3D: Generating 3D Adversarial Examples for 3D Object Detection in Driving Scenarios with NeRF
by: Li, Leheng, et al.
Published: (2023) -
SyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIs
by: Li, Leheng, et al.
Published: (2024)