DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yibo, Gao, Ruiyuan, Chen, Kai, Zhou, Kaiqiang, Cai, Yingjie, Hong, Lanqing, Li, Zhenguo, Jiang, Lihui, Yeung, Dit-Yan, Xu, Qiang, Zhang, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023)
by: Gao, Ruiyuan, et al.
Published: (2023)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
by: Gao, Ruiyuan, et al.
Published: (2024)
by: Gao, Ruiyuan, et al.
Published: (2024)
Mixed Autoencoder for Self-supervised Visual Representation Learning
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
Implicit Concept Removal of Diffusion Models
by: Liu, Zhili, et al.
Published: (2023)
by: Liu, Zhili, et al.
Published: (2023)
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
by: Gou, Yunhao, et al.
Published: (2024)
by: Gou, Yunhao, et al.
Published: (2024)
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
by: Chen, Kai, et al.
Published: (2024)
by: Chen, Kai, et al.
Published: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
CoherenDream: Boosting Holistic Text Coherence in 3D Generation via Multimodal Large Language Models Feedback
by: Jiang, Chenhan, et al.
Published: (2025)
by: Jiang, Chenhan, et al.
Published: (2025)
TransformMix: Learning Transformation and Mixing Strategies from Data
by: Cheung, Tsz-Him, et al.
Published: (2024)
by: Cheung, Tsz-Him, et al.
Published: (2024)
Mixture of Cluster-conditional LoRA Experts for Vision-language Instruction Tuning
by: Gou, Yunhao, et al.
Published: (2023)
by: Gou, Yunhao, et al.
Published: (2023)
IRWE: Inductive Random Walk for Joint Inference of Identity and Position Network Embedding
by: Qin, Meng, et al.
Published: (2024)
by: Qin, Meng, et al.
Published: (2024)
Unified Triplet-Level Hallucination Evaluation for Large Vision-Language Models
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake Analysis
by: Chen, Kai, et al.
Published: (2023)
by: Chen, Kai, et al.
Published: (2023)
Micro-Expression Recognition via Fine-Grained Dynamic Perception
by: Shao, Zhiwen, et al.
Published: (2025)
by: Shao, Zhiwen, et al.
Published: (2025)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
by: Zhao, Yuyang, et al.
Published: (2023)
by: Zhao, Yuyang, et al.
Published: (2023)
Non-Cross Diffusion for Semantic Consistency
by: Zheng, Ziyang, et al.
Published: (2023)
by: Zheng, Ziyang, et al.
Published: (2023)
SQS: Enhancing Sparse Perception Models via Query-based Splatting in Autonomous Driving
by: Zhang, Haiming, et al.
Published: (2025)
by: Zhang, Haiming, et al.
Published: (2025)
CVT-xRF: Contrastive In-Voxel Transformer for 3D Consistent Radiance Fields from Sparse Inputs
by: Zhong, Yingji, et al.
Published: (2024)
by: Zhong, Yingji, et al.
Published: (2024)
DetCLIPv3: Towards Versatile Generative Open-vocabulary Object Detection
by: Yao, Lewei, et al.
Published: (2024)
by: Yao, Lewei, et al.
Published: (2024)
Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks
by: Zeng, Kai, et al.
Published: (2025)
by: Zeng, Kai, et al.
Published: (2025)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
by: Liu, Zhili, et al.
Published: (2024)
by: Liu, Zhili, et al.
Published: (2024)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
by: Jiang, Chenhan, et al.
Published: (2026)
by: Jiang, Chenhan, et al.
Published: (2026)
JointDreamer: Ensuring Geometry Consistency and Text Congruence in Text-to-3D Generation via Joint Score Distillation
by: Jiang, Chenhan, et al.
Published: (2024)
by: Jiang, Chenhan, et al.
Published: (2024)
Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
by: Dong, Shiyin, et al.
Published: (2024)
by: Dong, Shiyin, et al.
Published: (2024)
Generation is Required for Data-Efficient Perception
by: Brady, Jack, et al.
Published: (2025)
by: Brady, Jack, et al.
Published: (2025)
PRJ: Perception-Retrieval-Judgement for Generated Images
by: Fu, Qiang, et al.
Published: (2025)
by: Fu, Qiang, et al.
Published: (2025)
Pre-train and Refine: Towards Higher Efficiency in K-Agnostic Community Detection without Quality Degradation
by: Qin, Meng, et al.
Published: (2024)
by: Qin, Meng, et al.
Published: (2024)
Rethinking Targeted Adversarial Attacks For Neural Machine Translation
by: Wu, Junjie, et al.
Published: (2024)
by: Wu, Junjie, et al.
Published: (2024)
Understanding LLMs' Fluid Intelligence Deficiency: An Analysis of the ARC Task
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems
by: Kai, Ding, et al.
Published: (2024)
by: Kai, Ding, et al.
Published: (2024)
STLDM: Spatio-Temporal Latent Diffusion Model for Precipitation Nowcasting
by: Foo, Shi Quan, et al.
Published: (2025)
by: Foo, Shi Quan, et al.
Published: (2025)
Getting More Juice Out of Your Data: Hard Pair Refinement Enhances Visual-Language Models Without Extra Data
by: Wang, Haonan, et al.
Published: (2023)
by: Wang, Haonan, et al.
Published: (2023)
Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel Views
by: Zhong, Yingji, et al.
Published: (2025)
by: Zhong, Yingji, et al.
Published: (2025)
VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
by: Cao, Yang, et al.
Published: (2026)
by: Cao, Yang, et al.
Published: (2026)
CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
by: Qi, Yu, et al.
Published: (2025)
by: Qi, Yu, et al.
Published: (2025)
Similar Items
-
MagicDrive: Street View Generation with Diverse 3D Geometry Control
by: Gao, Ruiyuan, et al.
Published: (2023) -
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
by: Chen, Kai, et al.
Published: (2023) -
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023) -
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
by: Gao, Ruiyuan, et al.
Published: (2024) -
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes
by: Gao, Ruiyuan, et al.
Published: (2024)