HuAMR: A Hungarian AMR Parser and Dataset

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Barta, Botond, Hamerlik, Endre, Nyist, Milán Konor, Ács, Judit
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915176378269696
author Barta, Botond
Hamerlik, Endre
Nyist, Milán Konor
Ács, Judit
author_facet Barta, Botond
Hamerlik, Endre
Nyist, Milán Konor
Ács, Judit
contents We present HuAMR, the first Abstract Meaning Representation (AMR) dataset and a suite of large language model-based AMR parsers for Hungarian, targeting the scarcity of semantic resources for non-English languages. To create HuAMR, we employed Llama-3.1-70B to automatically generate silver-standard AMR annotations, which we then refined manually to ensure quality. Building on this dataset, we investigate how different model architectures - mT5 Large and Llama-3.2-1B - and fine-tuning strategies affect AMR parsing performance. While incorporating silver-standard AMRs from Llama-3.1-70B into the training data of smaller models does not consistently boost overall scores, our results show that these techniques effectively enhance parsing accuracy on Hungarian news data (the domain of HuAMR). We evaluate our parsers using Smatch scores and confirm the potential of HuAMR and our parsers for advancing semantic parsing research.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20552
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HuAMR: A Hungarian AMR Parser and Dataset
Barta, Botond
Hamerlik, Endre
Nyist, Milán Konor
Ács, Judit
Computation and Language
We present HuAMR, the first Abstract Meaning Representation (AMR) dataset and a suite of large language model-based AMR parsers for Hungarian, targeting the scarcity of semantic resources for non-English languages. To create HuAMR, we employed Llama-3.1-70B to automatically generate silver-standard AMR annotations, which we then refined manually to ensure quality. Building on this dataset, we investigate how different model architectures - mT5 Large and Llama-3.2-1B - and fine-tuning strategies affect AMR parsing performance. While incorporating silver-standard AMRs from Llama-3.1-70B into the training data of smaller models does not consistently boost overall scores, our results show that these techniques effectively enhance parsing accuracy on Hungarian news data (the domain of HuAMR). We evaluate our parsers using Smatch scores and confirm the potential of HuAMR and our parsers for advancing semantic parsing research.
title HuAMR: A Hungarian AMR Parser and Dataset
topic Computation and Language
url https://arxiv.org/abs/2502.20552