Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adel, Mohamed, Alhafni, Bashar, Habash, Nizar
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910056381939712
author Adel, Mohamed
Alhafni, Bashar
Habash, Nizar
author_facet Adel, Mohamed
Alhafni, Bashar
Habash, Nizar
contents Large language models (LLMs) perform strongly on many NLP tasks, but their ability to produce explicit linguistic structure remains unclear. We evaluate instruction-tuned LLMs on two structured prediction tasks for Standard Arabic: morphosyntactic tagging and labeled dependency parsing. Arabic provides a challenging testbed due to its rich morphology and orthographic ambiguity, which create strong morphology-syntax interactions. We compare zero-shot prompting with retrieval-based in-context learning (ICL) using examples from Arabic treebanks. Results show that prompt design and demonstration selection strongly affect performance: proprietary models approach supervised baselines for feature-level tagging and become competitive with specialized dependency parsers. In raw-text settings, tokenization remains challenging, though retrieval-based ICL improves both parsing and tokenization. Our analysis highlights which aspects of Arabic morphosyntax and syntax LLMs capture reliably and which remain difficult.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16718
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models
Adel, Mohamed
Alhafni, Bashar
Habash, Nizar
Computation and Language
Large language models (LLMs) perform strongly on many NLP tasks, but their ability to produce explicit linguistic structure remains unclear. We evaluate instruction-tuned LLMs on two structured prediction tasks for Standard Arabic: morphosyntactic tagging and labeled dependency parsing. Arabic provides a challenging testbed due to its rich morphology and orthographic ambiguity, which create strong morphology-syntax interactions. We compare zero-shot prompting with retrieval-based in-context learning (ICL) using examples from Arabic treebanks. Results show that prompt design and demonstration selection strongly affect performance: proprietary models approach supervised baselines for feature-level tagging and become competitive with specialized dependency parsers. In raw-text settings, tokenization remains challenging, though retrieval-based ICL improves both parsing and tokenization. Our analysis highlights which aspects of Arabic morphosyntax and syntax LLMs capture reliably and which remain difficult.
title Arabic Morphosyntactic Tagging and Dependency Parsing with Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2603.16718