DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ahmad, Muhammad, Mazzara, Manuel, Distefano, Salvatore, Khan, Adil Mehmood, Ullo, Silvia Liberata
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909577781444608
author Ahmad, Muhammad
Mazzara, Manuel
Distefano, Salvatore
Khan, Adil Mehmood
Ullo, Silvia Liberata
author_facet Ahmad, Muhammad
Mazzara, Manuel
Distefano, Salvatore
Khan, Adil Mehmood
Ullo, Silvia Liberata
contents Hyperspectral image classification (HSIC) has gained significant attention because of its potential in analyzing high-dimensional data with rich spectral and spatial information. In this work, we propose the Differential Spatial-Spectral Transformer (DiffFormer), a novel framework designed to address the inherent challenges of HSIC, such as spectral redundancy and spatial discontinuity. The DiffFormer leverages a Differential Multi-Head Self-Attention (DMHSA) mechanism, which enhances local feature discrimination by introducing differential attention to accentuate subtle variations across neighboring spectral-spatial patches. The architecture integrates Spectral-Spatial Tokenization through three-dimensional (3D) convolution-based patch embeddings, positional encoding, and a stack of transformer layers equipped with the SWiGLU activation function for efficient feature extraction (SwiGLU is a variant of the Gated Linear Unit (GLU) activation function). A token-based classification head further ensures robust representation learning, enabling precise labeling of hyperspectral pixels. Extensive experiments on benchmark hyperspectral datasets demonstrate the superiority of DiffFormer in terms of classification accuracy, computational efficiency, and generalizability, compared to existing state-of-the-art (SOTA) methods. In addition, this work provides a detailed analysis of computational complexity, showcasing the scalability of the model for large-scale remote sensing applications. The source code will be made available at \url{https://github.com/mahmad000/DiffFormer} after the first round of revision.
format Preprint
id arxiv_https___arxiv_org_abs_2412_17350
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification
Ahmad, Muhammad
Mazzara, Manuel
Distefano, Salvatore
Khan, Adil Mehmood
Ullo, Silvia Liberata
Computer Vision and Pattern Recognition
Hyperspectral image classification (HSIC) has gained significant attention because of its potential in analyzing high-dimensional data with rich spectral and spatial information. In this work, we propose the Differential Spatial-Spectral Transformer (DiffFormer), a novel framework designed to address the inherent challenges of HSIC, such as spectral redundancy and spatial discontinuity. The DiffFormer leverages a Differential Multi-Head Self-Attention (DMHSA) mechanism, which enhances local feature discrimination by introducing differential attention to accentuate subtle variations across neighboring spectral-spatial patches. The architecture integrates Spectral-Spatial Tokenization through three-dimensional (3D) convolution-based patch embeddings, positional encoding, and a stack of transformer layers equipped with the SWiGLU activation function for efficient feature extraction (SwiGLU is a variant of the Gated Linear Unit (GLU) activation function). A token-based classification head further ensures robust representation learning, enabling precise labeling of hyperspectral pixels. Extensive experiments on benchmark hyperspectral datasets demonstrate the superiority of DiffFormer in terms of classification accuracy, computational efficiency, and generalizability, compared to existing state-of-the-art (SOTA) methods. In addition, this work provides a detailed analysis of computational complexity, showcasing the scalability of the model for large-scale remote sensing applications. The source code will be made available at \url{https://github.com/mahmad000/DiffFormer} after the first round of revision.
title DiffFormer: a Differential Spatial-Spectral Transformer for Hyperspectral Image Classification
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.17350