Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhuoran, Hu, Chunming, Chen, Junfan, Chen, Zhijun, Zhang, Richong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929759148048384
author Li, Zhuoran
Hu, Chunming
Chen, Junfan
Chen, Zhijun
Zhang, Richong
author_facet Li, Zhuoran
Hu, Chunming
Chen, Junfan
Chen, Zhijun
Zhang, Richong
contents Word order difference between source and target languages is a major obstacle to cross-lingual transfer, especially in the dependency parsing task. Current works are mostly based on order-agnostic models or word reordering to mitigate this problem. However, such methods either do not leverage grammatical information naturally contained in word order or are computationally expensive as the permutation space grows exponentially with the sentence length. Moreover, the reordered source sentence with an unnatural word order may be a form of noising that harms the model learning. To this end, we propose an Implicit Word Reordering framework with Knowledge Distillation (IWR-KD). This framework is inspired by that deep networks are good at learning feature linearization corresponding to meaningful data transformation, e.g. word reordering. To realize this idea, we introduce a knowledge distillation framework composed of a word-reordering teacher model and a dependency parsing student model. We verify our proposed method on Universal Dependency Treebanks across 31 different languages and show it outperforms a series of competitors, together with experimental analysis to illustrate how our method works towards training a robust parser.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17308
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
Li, Zhuoran
Hu, Chunming
Chen, Junfan
Chen, Zhijun
Zhang, Richong
Computation and Language
Machine Learning
Word order difference between source and target languages is a major obstacle to cross-lingual transfer, especially in the dependency parsing task. Current works are mostly based on order-agnostic models or word reordering to mitigate this problem. However, such methods either do not leverage grammatical information naturally contained in word order or are computationally expensive as the permutation space grows exponentially with the sentence length. Moreover, the reordered source sentence with an unnatural word order may be a form of noising that harms the model learning. To this end, we propose an Implicit Word Reordering framework with Knowledge Distillation (IWR-KD). This framework is inspired by that deep networks are good at learning feature linearization corresponding to meaningful data transformation, e.g. word reordering. To realize this idea, we introduce a knowledge distillation framework composed of a word-reordering teacher model and a dependency parsing student model. We verify our proposed method on Universal Dependency Treebanks across 31 different languages and show it outperforms a series of competitors, together with experimental analysis to illustrate how our method works towards training a robust parser.
title Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2502.17308