Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nigam, Shubham Kumar, Shukla, Parjanya Aditya, Shallum, Noel, Bhattacharya, Arnab
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908724629602304
author Nigam, Shubham Kumar
Shukla, Parjanya Aditya
Shallum, Noel
Bhattacharya, Arnab
author_facet Nigam, Shubham Kumar
Shukla, Parjanya Aditya
Shallum, Noel
Bhattacharya, Arnab
contents Handwritten text recognition (HTR) and machine translation continue to pose significant challenges, particularly for low-resource languages like Marathi, which lack large digitized corpora and exhibit high variability in handwriting styles. The conventional approach to address this involves a two-stage pipeline: an OCR system extracts text from handwritten images, which is then translated into the target language using a machine translation model. In this work, we explore and compare the performance of traditional OCR-MT pipelines with Vision Large Language Models that aim to unify these stages and directly translate handwritten text images in a single, end-to-end step. Our motivation is grounded in the urgent need for scalable, accurate translation systems to digitize legal records such as FIRs, charge sheets, and witness statements in India's district and high courts. We evaluate both approaches on a curated dataset of handwritten Marathi legal documents, with the goal of enabling efficient legal document processing, even in low-resource environments. Our findings offer actionable insights toward building robust, edge-deployable solutions that enhance access to legal information for non-native speakers and legal professionals alike.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
Nigam, Shubham Kumar
Shukla, Parjanya Aditya
Shallum, Noel
Bhattacharya, Arnab
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Handwritten text recognition (HTR) and machine translation continue to pose significant challenges, particularly for low-resource languages like Marathi, which lack large digitized corpora and exhibit high variability in handwriting styles. The conventional approach to address this involves a two-stage pipeline: an OCR system extracts text from handwritten images, which is then translated into the target language using a machine translation model. In this work, we explore and compare the performance of traditional OCR-MT pipelines with Vision Large Language Models that aim to unify these stages and directly translate handwritten text images in a single, end-to-end step. Our motivation is grounded in the urgent need for scalable, accurate translation systems to digitize legal records such as FIRs, charge sheets, and witness statements in India's district and high courts. We evaluate both approaches on a curated dataset of handwritten Marathi legal documents, with the goal of enabling efficient legal document processing, even in low-resource environments. Our findings offer actionable insights toward building robust, edge-deployable solutions that enhance access to legal information for non-native speakers and legal professionals alike.
title Seeing Justice Clearly: Handwritten Legal Document Translation with OCR and Vision-Language Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2512.18004