Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Reid, David, Arandjelovic, Ognjen
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915729699241984
author Reid, David
Arandjelovic, Ognjen
author_facet Reid, David
Arandjelovic, Ognjen
contents Automated analysis of ancient coins has the potential to help researchers extract more historical insights from large collections of coins and to help collectors understand what they are buying or selling. Recent research in this area has shown promise in focusing on identification of semantic elements as they are commonly depicted on ancient coins, by using convolutional neural networks (CNNs). This paper is the first to apply the recently proposed Vision Transformer (ViT) deep learning architecture to the task of identification of semantic elements on coins, using fully automatic learning from multi-modal data (images and unstructured text). This article summarises previous research in the area, discusses the training and implementation of ViT and CNN models for ancient coins analysis and provides an evaluation of their performance. The ViT models were found to outperform the newly trained CNN models in accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2601_09433
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?
Reid, David
Arandjelovic, Ognjen
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Automated analysis of ancient coins has the potential to help researchers extract more historical insights from large collections of coins and to help collectors understand what they are buying or selling. Recent research in this area has shown promise in focusing on identification of semantic elements as they are commonly depicted on ancient coins, by using convolutional neural networks (CNNs). This paper is the first to apply the recently proposed Vision Transformer (ViT) deep learning architecture to the task of identification of semantic elements on coins, using fully automatic learning from multi-modal data (images and unstructured text). This article summarises previous research in the area, discusses the training and implementation of ViT and CNN models for ancient coins analysis and provides an evaluation of their performance. The ViT models were found to outperform the newly trained CNN models in accuracy.
title Do Transformers Understand Ancient Roman Coin Motifs Better than CNNs?
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2601.09433