A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ye, Weixin, Wang, Wei, Liu, Yahui, Song, Yue, Ren, Bin, Bi, Wei, Cucchiara, Rita, Sebe, Nicu
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909993374056448
author Ye, Weixin
Wang, Wei
Liu, Yahui
Song, Yue
Ren, Bin
Bi, Wei
Cucchiara, Rita
Sebe, Nicu
author_facet Ye, Weixin
Wang, Wei
Liu, Yahui
Song, Yue
Ren, Bin
Bi, Wei
Cucchiara, Rita
Sebe, Nicu
contents In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable \textit{unknown (unk)} position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (\textit{e.g.,} ImageNet-1K) and sentiment analysis for text (\textit{e.g.,} Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks. Code is publicly available via https://github.com/ywxsuperstar/transformerattack
format Preprint
id arxiv_https___arxiv_org_abs_2601_12051
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
Ye, Weixin
Wang, Wei
Liu, Yahui
Song, Yue
Ren, Bin
Bi, Wei
Cucchiara, Rita
Sebe, Nicu
Computer Vision and Pattern Recognition
In federated learning, Transformer, as a popular architecture, faces critical challenges in defending against gradient attacks and improving model performance in both Computer Vision (CV) and Natural Language Processing (NLP) tasks. It has been revealed that the gradient of Position Embeddings (PEs) in Transformer contains sufficient information, which can be used to reconstruct the input data. To mitigate this issue, we introduce a Masked Jigsaw Puzzle (MJP) framework. MJP starts with random token shuffling to break the token order, and then a learnable \textit{unknown (unk)} position embedding is used to mask out the PEs of the shuffled tokens. In this manner, the local spatial information which is encoded in the position embeddings is disrupted, and the models are forced to learn feature representations that are less reliant on the local spatial information. Notably, with the careful use of MJP, we can not only improve models' robustness against gradient attacks, but also boost their performance in both vision and text application scenarios, such as classification for images (\textit{e.g.,} ImageNet-1K) and sentiment analysis for text (\textit{e.g.,} Yelp and Amazon). Experimental results suggest that MJP is a unified framework for different Transformer-based models in both vision and language tasks. Code is publicly available via https://github.com/ywxsuperstar/transformerattack
title A Unified Masked Jigsaw Puzzle Framework for Vision and Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2601.12051