Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Ziqi, Xu, Xin, Wang, Yu-Xiong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
by: Dong, Shiyin, et al.
Published: (2024)
by: Dong, Shiyin, et al.
Published: (2024)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
by: Pang, Ziqi, et al.
Published: (2025)
by: Pang, Ziqi, et al.
Published: (2025)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
by: Yin, Wen, et al.
Published: (2025)
by: Yin, Wen, et al.
Published: (2025)
Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator
by: Li, Jianze, et al.
Published: (2024)
by: Li, Jianze, et al.
Published: (2024)
Gait Recognition via Collaborating Discriminative and Generative Diffusion Models
by: Xiong, Haijun, et al.
Published: (2025)
by: Xiong, Haijun, et al.
Published: (2025)
SD-DiT: Unleashing the Power of Self-supervised Discrimination in Diffusion Transformer
by: Zhu, Rui, et al.
Published: (2024)
by: Zhu, Rui, et al.
Published: (2024)
Implicit and Explicit Language Guidance for Diffusion-based Visual Perception
by: Wang, Hefeng, et al.
Published: (2024)
by: Wang, Hefeng, et al.
Published: (2024)
AlignTok: Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
by: Chen, Bowei, et al.
Published: (2025)
by: Chen, Bowei, et al.
Published: (2025)
DanceGRPO: Unleashing GRPO on Visual Generation
by: Xue, Zeyue, et al.
Published: (2025)
by: Xue, Zeyue, et al.
Published: (2025)
GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation
by: Lin, Lang, et al.
Published: (2025)
by: Lin, Lang, et al.
Published: (2025)
Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
by: Gan, Chaofan, et al.
Published: (2025)
by: Gan, Chaofan, et al.
Published: (2025)
RMem: Restricted Memory Banks Improve Video Object Segmentation
by: Zhou, Junbao, et al.
Published: (2024)
by: Zhou, Junbao, et al.
Published: (2024)
Aligning Human Motion Generation with Human Perceptions
by: Wang, Haoru, et al.
Published: (2024)
by: Wang, Haoru, et al.
Published: (2024)
Illusions in Humans and AI: How Visual Perception Aligns and Diverges
by: Yang, Jianyi, et al.
Published: (2025)
by: Yang, Jianyi, et al.
Published: (2025)
Frozen Transformers in Language Models Are Effective Visual Encoder Layers
by: Pang, Ziqi, et al.
Published: (2023)
by: Pang, Ziqi, et al.
Published: (2023)
Unleashing Diffusion and State Space Models for Medical Image Segmentation
by: Wu, Rong, et al.
Published: (2025)
by: Wu, Rong, et al.
Published: (2025)
Denoise and Align: Diffusion-Driven Foreground Knowledge Prompting for Open-Vocabulary Temporal Action Detection
by: Zhu, Sa, et al.
Published: (2026)
by: Zhu, Sa, et al.
Published: (2026)
Brick-Diffusion: Generating Long Videos with Brick-to-Wall Denoising
by: Yuan, Yunlong, et al.
Published: (2025)
by: Yuan, Yunlong, et al.
Published: (2025)
Aligning Knowledge Graph with Visual Perception for Object-goal Navigation
by: Xu, Nuo, et al.
Published: (2024)
by: Xu, Nuo, et al.
Published: (2024)
Visual Bridge: Universal Visual Perception Representations Generating
by: Gao, Yilin, et al.
Published: (2025)
by: Gao, Yilin, et al.
Published: (2025)
RandAR: Decoder-only Autoregressive Visual Generation in Random Orders
by: Pang, Ziqi, et al.
Published: (2024)
by: Pang, Ziqi, et al.
Published: (2024)
Unleashing Perception-Time Scaling to Multimodal Reasoning Models
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Diff-Restorer: Unleashing Visual Prompts for Diffusion-based Universal Image Restoration
by: Zhang, Yuhong, et al.
Published: (2024)
by: Zhang, Yuhong, et al.
Published: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
by: Sun, Zhengyang, et al.
Published: (2026)
by: Sun, Zhengyang, et al.
Published: (2026)
Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
by: Wang, Ziyin, et al.
Published: (2026)
by: Wang, Ziyin, et al.
Published: (2026)
Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
by: Wang, Ziqi, et al.
Published: (2025)
by: Wang, Ziqi, et al.
Published: (2025)
G3PT: Unleash the power of Autoregressive Modeling in 3D Generation via Cross-scale Querying Transformer
by: Zhang, Jinzhi, et al.
Published: (2024)
by: Zhang, Jinzhi, et al.
Published: (2024)
Stimulating Diffusion Model for Image Denoising via Adaptive Embedding and Ensembling
by: Li, Tong, et al.
Published: (2023)
by: Li, Tong, et al.
Published: (2023)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
Step-level Denoising-time Diffusion Alignment with Multiple Objectives
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
PPTArena: A Benchmark for Agentic PowerPoint Editing
by: Ofengenden, Michael, et al.
Published: (2025)
by: Ofengenden, Michael, et al.
Published: (2025)
Step Saver: Predicting Minimum Denoising Steps for Diffusion Model Image Generation
by: Yu, Jean, et al.
Published: (2024)
by: Yu, Jean, et al.
Published: (2024)
Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
by: Gou, Yunhao, et al.
Published: (2025)
by: Gou, Yunhao, et al.
Published: (2025)
DetDiffusion: Synergizing Generative and Perceptive Models for Enhanced Data Generation and Perception
by: Wang, Yibo, et al.
Published: (2024)
by: Wang, Yibo, et al.
Published: (2024)
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Human-Aligned Generative Perception: Bridging Psychophysics and Generative Models
by: Titikhsha, Antara, et al.
Published: (2025)
by: Titikhsha, Antara, et al.
Published: (2025)
VFM-Loc: Zero-Shot Cross-View Geo-Localization via Aligning Discriminative Visual Hierarchies
by: Lu, Jun, et al.
Published: (2026)
by: Lu, Jun, et al.
Published: (2026)
Directly Denoising Diffusion Models
by: Zhang, Dan, et al.
Published: (2024)
by: Zhang, Dan, et al.
Published: (2024)
Adaptive and Iterative Point Cloud Denoising with Score-Based Diffusion Model
by: Wang, Zhaonan, et al.
Published: (2025)
by: Wang, Zhaonan, et al.
Published: (2025)
Similar Items
-
Bridging Generative and Discriminative Models for Unified Visual Perception with Diffusion Priors
by: Dong, Shiyin, et al.
Published: (2024) -
MR. Video: "MapReduce" is the Principle for Long Video Understanding
by: Pang, Ziqi, et al.
Published: (2025) -
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025) -
Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
by: Yin, Wen, et al.
Published: (2025) -
Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator
by: Li, Jianze, et al.
Published: (2024)