Saved in:
Bibliographic Details
Main Authors: Fort, Stanislav, Whitaker, Jonathan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2502.07753
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929710488879104
author Fort, Stanislav
Whitaker, Jonathan
author_facet Fort, Stanislav
Whitaker, Jonathan
contents We demonstrate that discriminative models inherently contain powerful generative capabilities, challenging the fundamental distinction between discriminative and generative architectures. Our method, Direct Ascent Synthesis (DAS), reveals these latent capabilities through multi-resolution optimization of CLIP model representations. While traditional inversion attempts produce adversarial patterns, DAS achieves high-quality image synthesis by decomposing optimization across multiple spatial scales (1x1 to 224x224), requiring no additional training. This approach not only enables diverse applications -- from text-to-image generation to style transfer -- but maintains natural image statistics ($1/f^2$ spectrum) and guides the generation away from non-robust adversarial patterns. Our results demonstrate that standard discriminative models encode substantially richer generative knowledge than previously recognized, providing new perspectives on model interpretability and the relationship between adversarial examples and natural image synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2502_07753
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
Fort, Stanislav
Whitaker, Jonathan
Computer Vision and Pattern Recognition
We demonstrate that discriminative models inherently contain powerful generative capabilities, challenging the fundamental distinction between discriminative and generative architectures. Our method, Direct Ascent Synthesis (DAS), reveals these latent capabilities through multi-resolution optimization of CLIP model representations. While traditional inversion attempts produce adversarial patterns, DAS achieves high-quality image synthesis by decomposing optimization across multiple spatial scales (1x1 to 224x224), requiring no additional training. This approach not only enables diverse applications -- from text-to-image generation to style transfer -- but maintains natural image statistics ($1/f^2$ spectrum) and guides the generation away from non-robust adversarial patterns. Our results demonstrate that standard discriminative models encode substantially richer generative knowledge than previously recognized, providing new perspectives on model interpretability and the relationship between adversarial examples and natural image synthesis.
title Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2502.07753