Cross-domain Chinese Sentence Pattern Parsing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yu, Jingsi, Kong, Cunliang, Yang, Liner, Zhang, Meishan, Zhu, Lin, Wang, Yujie, Lin, Haozhe, Sun, Maosong, Yang, Erhong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914743250321408
author Yu, Jingsi
Kong, Cunliang
Yang, Liner
Zhang, Meishan
Zhu, Lin
Wang, Yujie
Lin, Haozhe
Sun, Maosong
Yang, Erhong
author_facet Yu, Jingsi
Kong, Cunliang
Yang, Liner
Zhang, Meishan
Zhu, Lin
Wang, Yujie
Lin, Haozhe
Sun, Maosong
Yang, Erhong
contents Sentence Pattern Structure (SPS) parsing is a syntactic analysis method primarily employed in language teaching.Existing SPS parsers rely heavily on textbook corpora for training, lacking cross-domain capability.To overcome this constraint, this paper proposes an innovative approach leveraging large language models (LLMs) within a self-training framework. Partial syntactic rules from a source domain are combined with target domain sentences to dynamically generate training data, enhancing the adaptability of the parser to diverse domains.Experiments conducted on textbook and news domains demonstrate the effectiveness of the proposed method, outperforming rule-based baselines by 1.68 points on F1 metrics.
format Preprint
id arxiv_https___arxiv_org_abs_2402_16311
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-domain Chinese Sentence Pattern Parsing
Yu, Jingsi
Kong, Cunliang
Yang, Liner
Zhang, Meishan
Zhu, Lin
Wang, Yujie
Lin, Haozhe
Sun, Maosong
Yang, Erhong
Computation and Language
Artificial Intelligence
Sentence Pattern Structure (SPS) parsing is a syntactic analysis method primarily employed in language teaching.Existing SPS parsers rely heavily on textbook corpora for training, lacking cross-domain capability.To overcome this constraint, this paper proposes an innovative approach leveraging large language models (LLMs) within a self-training framework. Partial syntactic rules from a source domain are combined with target domain sentences to dynamically generate training data, enhancing the adaptability of the parser to diverse domains.Experiments conducted on textbook and news domains demonstrate the effectiveness of the proposed method, outperforming rule-based baselines by 1.68 points on F1 metrics.
title Cross-domain Chinese Sentence Pattern Parsing
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2402.16311