Inverting Neural Networks: New Methods to Generate Neural Network Inputs from Prescribed Outputs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pattichis, Rebecca, Janampa, Sebastian, Pattichis, Constantinos S., Pattichis, Marios S.
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915885634027520
author Pattichis, Rebecca
Janampa, Sebastian
Pattichis, Constantinos S.
Pattichis, Marios S.
author_facet Pattichis, Rebecca
Janampa, Sebastian
Pattichis, Constantinos S.
Pattichis, Marios S.
contents Neural network systems describe complex mappings that can be very difficult to understand. In this paper, we study the inverse problem of determining the input images that get mapped to specific neural network classes. Ultimately, we expect that these images contain recognizable features that are associated with their corresponding class classifications. We introduce two general methods for solving the inverse problem. In our forward pass method, we develop an inverse method based on a root-finding algorithm and the Jacobian with respect to the input image. In our backward pass method, we iteratively invert each layer, at the top. During the inversion process, we add random vectors sampled from the null-space of each linear layer. We demonstrate our new methods on both transformer architectures and sequential networks based on linear layers. Unlike previous methods, we show that our new methods are able to produce random-like input images that yield near perfect classification scores in all cases, revealing vulnerabilities in the underlying networks. Hence, we conclude that the proposed methods provide a more comprehensive coverage of the input image spaces that solve the inverse mapping problem.
format Preprint
id arxiv_https___arxiv_org_abs_2603_20461
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Inverting Neural Networks: New Methods to Generate Neural Network Inputs from Prescribed Outputs
Pattichis, Rebecca
Janampa, Sebastian
Pattichis, Constantinos S.
Pattichis, Marios S.
Computer Vision and Pattern Recognition
Neural network systems describe complex mappings that can be very difficult to understand. In this paper, we study the inverse problem of determining the input images that get mapped to specific neural network classes. Ultimately, we expect that these images contain recognizable features that are associated with their corresponding class classifications. We introduce two general methods for solving the inverse problem. In our forward pass method, we develop an inverse method based on a root-finding algorithm and the Jacobian with respect to the input image. In our backward pass method, we iteratively invert each layer, at the top. During the inversion process, we add random vectors sampled from the null-space of each linear layer. We demonstrate our new methods on both transformer architectures and sequential networks based on linear layers. Unlike previous methods, we show that our new methods are able to produce random-like input images that yield near perfect classification scores in all cases, revealing vulnerabilities in the underlying networks. Hence, we conclude that the proposed methods provide a more comprehensive coverage of the input image spaces that solve the inverse mapping problem.
title Inverting Neural Networks: New Methods to Generate Neural Network Inputs from Prescribed Outputs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.20461