Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhu, Xuekang, Ma, Xiaochen, Su, Lei, Jiang, Zhuohang, Du, Bo, Wang, Xiwen, Lei, Zeyu, Feng, Wentao, Pun, Chi-Man, Zhou, Jizhe
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866929637570904064
author Zhu, Xuekang
Ma, Xiaochen
Su, Lei
Jiang, Zhuohang
Du, Bo
Wang, Xiwen
Lei, Zeyu
Feng, Wentao
Pun, Chi-Man
Zhou, Jizhe
author_facet Zhu, Xuekang
Ma, Xiaochen
Su, Lei
Jiang, Zhuohang
Du, Bo
Wang, Xiwen
Lei, Zeyu
Feng, Wentao
Pun, Chi-Man
Zhou, Jizhe
contents The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13753
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization
Zhu, Xuekang
Ma, Xiaochen
Su, Lei
Jiang, Zhuohang
Du, Bo
Wang, Xiwen
Lei, Zeyu
Feng, Wentao
Pun, Chi-Man
Zhou, Jizhe
Computer Vision and Pattern Recognition
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
title Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.13753