Learning to Decipher from Pixels -- A Case Study of Copiale

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kang, Lei, De Gregorio, Giuseppe, Heil, Raphaela, Fornés, Alicia, Megyesi, Beáta
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913063180959744
author Kang, Lei
De Gregorio, Giuseppe
Heil, Raphaela
Fornés, Alicia
Megyesi, Beáta
author_facet Kang, Lei
De Gregorio, Giuseppe
Heil, Raphaela
Fornés, Alicia
Megyesi, Beáta
contents Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription-first paradigm, in which handwritten symbols are transcribed prior to decipherment. This intermediate step is labor-intensive, error-prone, and not always aligned with the goal of direct plaintext recovery. We propose an end-to-end, transcription-free approach that directly maps handwritten cipher images to plaintext. Using the Copiale cipher as a case study, we introduce the first text-line-level dataset pairing cipher images with German plaintext. We show that pretraining on generic handwriting data followed by cipher-specific fine-tuning substantially improves decipherment accuracy. Our results demonstrate that transcription-free image-to-plaintext decipherment is both feasible and effective for historical substitution ciphers, offering a simplified and scalable alternative to traditional pipelines. https://github.com/leitro/Decipher-from-Pixels-Copiale
format Preprint
id arxiv_https___arxiv_org_abs_2604_23683
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Decipher from Pixels -- A Case Study of Copiale
Kang, Lei
De Gregorio, Giuseppe
Heil, Raphaela
Fornés, Alicia
Megyesi, Beáta
Computer Vision and Pattern Recognition
Historical encrypted manuscripts require both paleographic interpretation of cipher symbols and cryptanalytic recovery of plaintext. Most existing computational workflows rely on a transcription-first paradigm, in which handwritten symbols are transcribed prior to decipherment. This intermediate step is labor-intensive, error-prone, and not always aligned with the goal of direct plaintext recovery. We propose an end-to-end, transcription-free approach that directly maps handwritten cipher images to plaintext. Using the Copiale cipher as a case study, we introduce the first text-line-level dataset pairing cipher images with German plaintext. We show that pretraining on generic handwriting data followed by cipher-specific fine-tuning substantially improves decipherment accuracy. Our results demonstrate that transcription-free image-to-plaintext decipherment is both feasible and effective for historical substitution ciphers, offering a simplified and scalable alternative to traditional pipelines. https://github.com/leitro/Decipher-from-Pixels-Copiale
title Learning to Decipher from Pixels -- A Case Study of Copiale
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.23683