Second language Korean Universal Dependency treebank v1.2: Focus on data augmentation and annotation scheme refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Sung, Hakyung, Shin, Gyu-Ho |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation
by: Sung, Hakyung, et al.
Published: (2026)
by: Sung, Hakyung, et al.
Published: (2026)
UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags
by: Sung, Hakyung, et al.
Published: (2025)
by: Sung, Hakyung, et al.
Published: (2025)
ASC analyzer: A Python package for measuring argument structure construction usage in English texts
by: Sung, Hakyung, et al.
Published: (2025)
by: Sung, Hakyung, et al.
Published: (2025)
Comparing human and LLM proofreading in L2 writing: Impact on lexical and syntactic features
by: Sung, Hakyung, et al.
Published: (2025)
by: Sung, Hakyung, et al.
Published: (2025)
Punctuation-aware treebank tree binarization
by: Klinger, Eitan, et al.
Published: (2025)
by: Klinger, Eitan, et al.
Published: (2025)
Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
by: Dobrovoljc, Kaja
Published: (2025)
by: Dobrovoljc, Kaja
Published: (2025)
K-UD: Revising Korean Universal Dependencies Guidelines
by: Kim, Kyuwon, et al.
Published: (2024)
by: Kim, Kyuwon, et al.
Published: (2024)
Construction and educational application of a linguistically grounded dependency treebank for Uyghur
by: Zuo, Jiaxin, et al.
Published: (2025)
by: Zuo, Jiaxin, et al.
Published: (2025)
The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Do language models capture implied discourse meanings? An investigation with exhaustivity implicatures of Korean morphology
by: Shin, Hagyeong, et al.
Published: (2024)
by: Shin, Hagyeong, et al.
Published: (2024)
Does Incomplete Syntax Influence Korean Language Model? Focusing on Word Order and Case Markers
by: Kim, Jong Myoung, et al.
Published: (2024)
by: Kim, Jong Myoung, et al.
Published: (2024)
Retrieval augmentation of large language models for lay language generation
by: Guo, Yue, et al.
Published: (2022)
by: Guo, Yue, et al.
Published: (2022)
KOMBO: Korean Character Representations Based on the Combination Rules of Subcharacters
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
Dravidian language family through Universal Dependencies lens
by: Rama, Taraka, et al.
Published: (2024)
by: Rama, Taraka, et al.
Published: (2024)
Enhancing Korean Dependency Parsing with Morphosyntactic Features
by: Park, Jungyeul, et al.
Published: (2025)
by: Park, Jungyeul, et al.
Published: (2025)
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
by: Kim, SungHo, et al.
Published: (2025)
by: Kim, SungHo, et al.
Published: (2025)
Large language models struggle with ethnographic text annotation
by: Goodall, Leonardo S., et al.
Published: (2026)
by: Goodall, Leonardo S., et al.
Published: (2026)
Exploring transfer learning for Deep NLP systems on rarely annotated languages
by: Yadav, Dipendra, et al.
Published: (2024)
by: Yadav, Dipendra, et al.
Published: (2024)
EVOKE: Emotion Vocabulary Of Korean and English
by: Jung, Yoonwon, et al.
Published: (2026)
by: Jung, Yoonwon, et al.
Published: (2026)
SCRIPT: A Subcharacter Compositional Representation Injection Module for Korean Pre-Trained Language Models
by: Kim, SungHo, et al.
Published: (2026)
by: Kim, SungHo, et al.
Published: (2026)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
KIT-19: A Comprehensive Korean Instruction Toolkit on 19 Tasks for Fine-Tuning Korean Large Language Models
by: Jang, Dongjun, et al.
Published: (2024)
by: Jang, Dongjun, et al.
Published: (2024)
Evaluating Large language models on Understanding Korean indirect Speech acts
by: Koo, Youngeun, et al.
Published: (2025)
by: Koo, Youngeun, et al.
Published: (2025)
Efficient argument classification with compact language models and ChatGPT-4 refinements
by: Pietron, Marcin, et al.
Published: (2024)
by: Pietron, Marcin, et al.
Published: (2024)
Universal Dependencies for the AnCora treebanks
by: Héctor Martínez Alonso
Published: (2016)
by: Héctor Martínez Alonso
Published: (2016)
CGELBank Annotation Manual v1.2
by: Reynolds, Brett, et al.
Published: (2023)
by: Reynolds, Brett, et al.
Published: (2023)
Large corpora and large language models: a replicable method for automating grammatical annotation
by: Morin, Cameron, et al.
Published: (2024)
by: Morin, Cameron, et al.
Published: (2024)
ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical language
by: Górski, Franciszek, et al.
Published: (2025)
by: Górski, Franciszek, et al.
Published: (2025)
Octopus v2: On-device language model for super agent
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
Tgea: An error-annotated dataset and benchmark tasks for text generation from pretrained language models
by: He, Jie, et al.
Published: (2025)
by: He, Jie, et al.
Published: (2025)
Octopus v4: Graph of language models
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
BoAT v2 -- A Web-Based Dependency Annotation Tool with Focus on Agglutinative Languages
by: Akkurt, Salih Furkan, et al.
Published: (2022)
by: Akkurt, Salih Furkan, et al.
Published: (2022)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
by: Lyth, Dan, et al.
Published: (2024)
by: Lyth, Dan, et al.
Published: (2024)
Is linguistically-motivated data augmentation worth it?
by: Groshan, Ray, et al.
Published: (2025)
by: Groshan, Ray, et al.
Published: (2025)
Improving accuracy of GPT-3/4 results on biomedical data using a retrieval-augmented language model
by: Soong, David, et al.
Published: (2023)
by: Soong, David, et al.
Published: (2023)
KnowsLM: A framework for evaluation of small language models for knowledge augmentation and humanised conversations
by: Harbola, Chitranshu, et al.
Published: (2025)
by: Harbola, Chitranshu, et al.
Published: (2025)
Coconstructions in spoken data: UD annotation guidelines and first results
by: Pannitto, Ludovica, et al.
Published: (2026)
by: Pannitto, Ludovica, et al.
Published: (2026)
KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
by: Wang, Xiaonan, et al.
Published: (2024)
by: Wang, Xiaonan, et al.
Published: (2024)
Similar Items
-
Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation
by: Sung, Hakyung, et al.
Published: (2026) -
UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags
by: Sung, Hakyung, et al.
Published: (2025) -
ASC analyzer: A Python package for measuring argument structure construction usage in English texts
by: Sung, Hakyung, et al.
Published: (2025) -
Comparing human and LLM proofreading in L2 writing: Impact on lexical and syntactic features
by: Sung, Hakyung, et al.
Published: (2025) -
Punctuation-aware treebank tree binarization
by: Klinger, Eitan, et al.
Published: (2025)