BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Huang, Tianyuan, Zhu, Zepeng, Xing, Hangdi, Shao, Zirui, Yu, Zhi, Yang, Chaoxiong, He, Jiaxian, Liu, Xiaozhong, Bu, Jiajun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915566276575232
author Huang, Tianyuan
Zhu, Zepeng
Xing, Hangdi
Shao, Zirui
Yu, Zhi
Yang, Chaoxiong
He, Jiaxian
Liu, Xiaozhong
Bu, Jiajun
author_facet Huang, Tianyuan
Zhu, Zepeng
Xing, Hangdi
Shao, Zirui
Yu, Zhi
Yang, Chaoxiong
He, Jiaxian
Liu, Xiaozhong
Bu, Jiajun
contents Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarcity and ambiguities in mixed-text contexts. We construct English and Chinese Braille Mixed Datasets (EBMD/CBMD) with mathematical formulas to support diverse Braille domain research, and propose a syntax tree-based augmentation method tailored for Braille data. To address the underperformance of traditional fine-tuning methods in Braille-related tasks, we investigate Braille Knowledge-Based Fine-Tuning (BKFT), which reduces the learning difficulty of Braille contextual features. BrailleLLM employs BKFT via instruction tuning to achieve unified Braille translation, formula-to-Braille conversion, and mixed-text translation. Experiments demonstrate that BKFT achieves significant performance improvements over conventional fine-tuning in Braille translation scenarios. Our open-sourced datasets and methodologies establish a foundation for low-resource multilingual Braille research.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18288
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
Huang, Tianyuan
Zhu, Zepeng
Xing, Hangdi
Shao, Zirui
Yu, Zhi
Yang, Chaoxiong
He, Jiaxian
Liu, Xiaozhong
Bu, Jiajun
Computation and Language
Braille plays a vital role in education and information accessibility for visually impaired individuals. However, Braille information processing faces challenges such as data scarcity and ambiguities in mixed-text contexts. We construct English and Chinese Braille Mixed Datasets (EBMD/CBMD) with mathematical formulas to support diverse Braille domain research, and propose a syntax tree-based augmentation method tailored for Braille data. To address the underperformance of traditional fine-tuning methods in Braille-related tasks, we investigate Braille Knowledge-Based Fine-Tuning (BKFT), which reduces the learning difficulty of Braille contextual features. BrailleLLM employs BKFT via instruction tuning to achieve unified Braille translation, formula-to-Braille conversion, and mixed-text translation. Experiments demonstrate that BKFT achieves significant performance improvements over conventional fine-tuning in Braille translation scenarios. Our open-sourced datasets and methodologies establish a foundation for low-resource multilingual Braille research.
title BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
topic Computation and Language
url https://arxiv.org/abs/2510.18288