Inferring the Graph Structure of Images for Graph Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gowda, Mayur S, Shi, John, Santos, Augusto, Moura, José M. F.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911138866790400
author Gowda, Mayur S
Shi, John
Santos, Augusto
Moura, José M. F.
author_facet Gowda, Mayur S
Shi, John
Santos, Augusto
Moura, José M. F.
contents Image datasets such as MNIST are a key benchmark for testing Graph Neural Network (GNN) architectures. The images are traditionally represented as a grid graph with each node representing a pixel and edges connecting neighboring pixels (vertically and horizontally). The graph signal is the values (intensities) of each pixel in the image. The graphs are commonly used as input to graph neural networks (e.g., Graph Convolutional Neural Networks (Graph CNNs) [1, 2], Graph Attention Networks (GAT) [3], GatedGCN [4]) to classify the images. In this work, we improve the accuracy of downstream graph neural network tasks by finding alternative graphs to the grid graph and superpixel methods to represent the dataset images, following the approach in [5, 6]. We find row correlation, column correlation, and product graphs for each image in MNIST and Fashion-MNIST using correlations between the pixel values building on the method in [5, 6]. Experiments show that using these different graph representations and features as input into downstream GNN models improves the accuracy over using the traditional grid graph and superpixel methods in the literature.
format Preprint
id arxiv_https___arxiv_org_abs_2509_04677
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Inferring the Graph Structure of Images for Graph Neural Networks
Gowda, Mayur S
Shi, John
Santos, Augusto
Moura, José M. F.
Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Signal Processing
Image datasets such as MNIST are a key benchmark for testing Graph Neural Network (GNN) architectures. The images are traditionally represented as a grid graph with each node representing a pixel and edges connecting neighboring pixels (vertically and horizontally). The graph signal is the values (intensities) of each pixel in the image. The graphs are commonly used as input to graph neural networks (e.g., Graph Convolutional Neural Networks (Graph CNNs) [1, 2], Graph Attention Networks (GAT) [3], GatedGCN [4]) to classify the images. In this work, we improve the accuracy of downstream graph neural network tasks by finding alternative graphs to the grid graph and superpixel methods to represent the dataset images, following the approach in [5, 6]. We find row correlation, column correlation, and product graphs for each image in MNIST and Fashion-MNIST using correlations between the pixel values building on the method in [5, 6]. Experiments show that using these different graph representations and features as input into downstream GNN models improves the accuracy over using the traditional grid graph and superpixel methods in the literature.
title Inferring the Graph Structure of Images for Graph Neural Networks
topic Image and Video Processing
Computer Vision and Pattern Recognition
Machine Learning
Signal Processing
url https://arxiv.org/abs/2509.04677