Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wan, Zihao, Xu, Pau Tong Lin, Luo, Fuwen, Wang, Ziyue, Li, Peng, Liu, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909880804179968
author Wan, Zihao
Xu, Pau Tong Lin
Luo, Fuwen
Wang, Ziyue
Li, Peng
Liu, Yang
author_facet Wan, Zihao
Xu, Pau Tong Lin
Luo, Fuwen
Wang, Ziyue
Li, Peng
Liu, Yang
contents While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less explored. Pictographic characters, which combine visual form with symbolic structure, provide an ideal test case for this capability. We formulate this visual recognition challenge in the mathematical domain, where each character is represented by an executable program of geometric primitives. This is framed as a program synthesis task, training a VLM to decompile raster images into programs composed of Bézier curves. Our model, acting as a "visual decompiler", demonstrates performance superior to strong zero-shot baselines, including GPT-4o. The most significant finding is that when trained solely on modern Chinese characters, the model is able to reconstruct ancient Oracle Bone Script in a zero-shot context. This generalization provides strong evidence that the model acquires an abstract and transferable geometric grammar, moving beyond pixel-level pattern recognition to a more structured form of visual understanding.
format Preprint
id arxiv_https___arxiv_org_abs_2511_00076
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
Wan, Zihao
Xu, Pau Tong Lin
Luo, Fuwen
Wang, Ziyue
Li, Peng
Liu, Yang
Machine Learning
While Vision-language Models (VLMs) have demonstrated strong semantic capabilities, their ability to interpret the underlying geometric structure of visual information is less explored. Pictographic characters, which combine visual form with symbolic structure, provide an ideal test case for this capability. We formulate this visual recognition challenge in the mathematical domain, where each character is represented by an executable program of geometric primitives. This is framed as a program synthesis task, training a VLM to decompile raster images into programs composed of Bézier curves. Our model, acting as a "visual decompiler", demonstrates performance superior to strong zero-shot baselines, including GPT-4o. The most significant finding is that when trained solely on modern Chinese characters, the model is able to reconstruct ancient Oracle Bone Script in a zero-shot context. This generalization provides strong evidence that the model acquires an abstract and transferable geometric grammar, moving beyond pixel-level pattern recognition to a more structured form of visual understanding.
title Bridging Vision, Language, and Mathematics: Pictographic Character Reconstruction with Bézier Curves
topic Machine Learning
url https://arxiv.org/abs/2511.00076