Toward Real Text Manipulation Detection: New Dataset and New Solution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Luo, Dongliang, Liu, Yuliang, Yang, Rui, Liu, Xianjin, Zeng, Jishen, Zhou, Yu, Bai, Xiang
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929219233120256
author Luo, Dongliang
Liu, Yuliang
Yang, Rui
Liu, Xianjin
Zeng, Jishen
Zhou, Yu
Bai, Xiang
author_facet Luo, Dongliang
Liu, Yuliang
Yang, Rui
Liu, Xianjin
Zeng, Jishen
Zhou, Yu
Bai, Xiang
contents With the surge in realistic text tampering, detecting fraudulent text in images has gained prominence for maintaining information security. However, the high costs associated with professional text manipulation and annotation limit the availability of real-world datasets, with most relying on synthetic tampering, which inadequately replicates real-world tampering attributes. To address this issue, we present the Real Text Manipulation (RTM) dataset, encompassing 14,250 text images, which include 5,986 manually and 5,258 automatically tampered images, created using a variety of techniques, alongside 3,006 unaltered text images for evaluating solution stability. Our evaluations indicate that existing methods falter in text forgery detection on the RTM dataset. We propose a robust baseline solution featuring a Consistency-aware Aggregation Hub and a Gated Cross Neighborhood-attention Fusion module for efficient multi-modal information fusion, supplemented by a Tampered-Authentic Contrastive Learning module during training, enriching feature representation distinction. This framework, extendable to other dual-stream architectures, demonstrated notable localization performance improvements of 7.33% and 6.38% on manual and overall manipulations, respectively. Our contributions aim to propel advancements in real-world text tampering detection. Code and dataset will be made available at https://github.com/DrLuo/RTM
format Preprint
id arxiv_https___arxiv_org_abs_2312_06934
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Toward Real Text Manipulation Detection: New Dataset and New Solution
Luo, Dongliang
Liu, Yuliang
Yang, Rui
Liu, Xianjin
Zeng, Jishen
Zhou, Yu
Bai, Xiang
Computer Vision and Pattern Recognition
With the surge in realistic text tampering, detecting fraudulent text in images has gained prominence for maintaining information security. However, the high costs associated with professional text manipulation and annotation limit the availability of real-world datasets, with most relying on synthetic tampering, which inadequately replicates real-world tampering attributes. To address this issue, we present the Real Text Manipulation (RTM) dataset, encompassing 14,250 text images, which include 5,986 manually and 5,258 automatically tampered images, created using a variety of techniques, alongside 3,006 unaltered text images for evaluating solution stability. Our evaluations indicate that existing methods falter in text forgery detection on the RTM dataset. We propose a robust baseline solution featuring a Consistency-aware Aggregation Hub and a Gated Cross Neighborhood-attention Fusion module for efficient multi-modal information fusion, supplemented by a Tampered-Authentic Contrastive Learning module during training, enriching feature representation distinction. This framework, extendable to other dual-stream architectures, demonstrated notable localization performance improvements of 7.33% and 6.38% on manual and overall manipulations, respectively. Our contributions aim to propel advancements in real-world text tampering detection. Code and dataset will be made available at https://github.com/DrLuo/RTM
title Toward Real Text Manipulation Detection: New Dataset and New Solution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.06934