Coconstructions in spoken data: UD annotation guidelines and first results
Fuente:
arXiv
Saved in:
| Main Authors: | Pannitto, Ludovica, Kahane, Sylvain, Dobrovoljc, Kaja, Battaglia, Elena, Guillaume, Bruno, Mauri, Caterina, Zucchini, Eleonora |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Towards the first UD Treebank of Spoken Italian: the KIParla forest
by: Pannitto, Ludovica
Published: (2024)
by: Pannitto, Ludovica
Published: (2024)
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
by: Simonotti, Martina, et al.
Published: (2026)
by: Simonotti, Martina, et al.
Published: (2026)
Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
by: Dobrovoljc, Kaja
Published: (2025)
by: Dobrovoljc, Kaja
Published: (2025)
Annotating Constructions with UD: the experience of the Italian Constructicon
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Recurrent babbling: evaluating the acquisition of grammar from limited input data
by: Pannitto, Ludovica, et al.
Published: (2020)
by: Pannitto, Ludovica, et al.
Published: (2020)
Linguistic Characteristics of AI-Generated Text: A Survey
by: Terčon, Luka, et al.
Published: (2025)
by: Terčon, Luka, et al.
Published: (2025)
CALaMo: a Constructionist Assessment of Language Models
by: Pannitto, Ludovica, et al.
Published: (2023)
by: Pannitto, Ludovica, et al.
Published: (2023)
Constraining constructions with WordNet: pros and cons for the semantic annotation of fillers in the Italian Constructicon
by: Pisciotta, Flavio, et al.
Published: (2025)
by: Pisciotta, Flavio, et al.
Published: (2025)
Did somebody say "Gest-IT"? A pilot exploration of multimodal data management
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
'Layer su Layer': Identifying and Disambiguating the Italian NPN Construction in BERT's family
by: Gorzoni, Greta, et al.
Published: (2026)
by: Gorzoni, Greta, et al.
Published: (2026)
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
by: Arčon, Tjaša, et al.
Published: (2026)
by: Arčon, Tjaša, et al.
Published: (2026)
Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
by: Pranjić, Marko, et al.
Published: (2024)
by: Pranjić, Marko, et al.
Published: (2024)
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
by: Klemen, Matej, et al.
Published: (2025)
by: Klemen, Matej, et al.
Published: (2025)
Sparse Logistic Regression with High-order Features for Automatic Grammar Rule Extraction from Treebanks
by: Herrera, Santiago, et al.
Published: (2024)
by: Herrera, Santiago, et al.
Published: (2024)
Parser agreement and disagreement in L2 Korean UD: Implications for human-in-the-loop annotation
by: Sung, Hakyung, et al.
Published: (2026)
by: Sung, Hakyung, et al.
Published: (2026)
Syntaxe théorique et formelle
by: Kahane, Sylvain, et al.
Published: (2023)
by: Kahane, Sylvain, et al.
Published: (2023)
A UD Treebank for Bohairic Coptic
by: Zeldes, Amir, et al.
Published: (2025)
by: Zeldes, Amir, et al.
Published: (2025)
Building UD Cairo for Old English in the Classroom
by: Levine, Lauren, et al.
Published: (2025)
by: Levine, Lauren, et al.
Published: (2025)
K-UD: Revising Korean Universal Dependencies Guidelines
by: Kim, Kyuwon, et al.
Published: (2024)
by: Kim, Kyuwon, et al.
Published: (2024)
Aligning the Norwegian UD Treebank with Entity and Coreference Information
by: Jørgensen, Tollef Emil, et al.
Published: (2023)
by: Jørgensen, Tollef Emil, et al.
Published: (2023)
Universal NER v2: Towards a Massively Multilingual Named Entity Recognition Benchmark
by: Blevins, Terra, et al.
Published: (2026)
by: Blevins, Terra, et al.
Published: (2026)
Out-of-distribution generalisation in spoken language understanding
by: Porjazovski, Dejan, et al.
Published: (2024)
by: Porjazovski, Dejan, et al.
Published: (2024)
Exploring Multiple Strategies to Improve Multilingual Coreference Resolution in CorefUD
by: Pražák, Ondřej, et al.
Published: (2024)
by: Pražák, Ondřej, et al.
Published: (2024)
Lost in Speech: Benchmarking, Evaluation, and Parsing of Spoken Code-Switching Beyond Standard UD Assumptions
by: Tyagi, Nemika, et al.
Published: (2026)
by: Tyagi, Nemika, et al.
Published: (2026)
Encoding of lexical tone in self-supervised models of spoken language
by: Shen, Gaofei, et al.
Published: (2024)
by: Shen, Gaofei, et al.
Published: (2024)
Empirical evidence of Large Language Model's influence on human spoken communication
by: Yakura, Hiromu, et al.
Published: (2024)
by: Yakura, Hiromu, et al.
Published: (2024)
Scaling few-shot spoken word classification with generative meta-continual learning
by: Beyers, Louise, et al.
Published: (2026)
by: Beyers, Louise, et al.
Published: (2026)
Extracting accent features in spoken Brazilian Portuguese without sociolinguistic labels
by: Leite, Pedro H. L., et al.
Published: (2026)
by: Leite, Pedro H. L., et al.
Published: (2026)
The UD-NewsCrawl Treebank: Reflections and Challenges from a Large-scale Tagalog Syntactic Annotation Project
by: Aquino, Angelina A., et al.
Published: (2025)
by: Aquino, Angelina A., et al.
Published: (2025)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
by: Faye, Géraud, et al.
Published: (2024)
by: Faye, Géraud, et al.
Published: (2024)
Tailoring AI-Driven Reading Scaffolds to the Distinct Needs of Neurodiverse Learners
by: Jhilal, Soufiane, et al.
Published: (2026)
by: Jhilal, Soufiane, et al.
Published: (2026)
The realization of tones in spontaneous spoken Taiwan Mandarin: a corpus-based survey and theory-driven computational modeling
by: Lu, Yuxin, et al.
Published: (2025)
by: Lu, Yuxin, et al.
Published: (2025)
Parsing the Switch: LLM-Based UD Annotation for Complex Code-Switched and Low-Resource Languages
by: Kellert, Olga, et al.
Published: (2025)
by: Kellert, Olga, et al.
Published: (2025)
LLM_annotate: A Python package for annotating and analyzing fiction characters
by: Rosenbusch, Hannes
Published: (2025)
by: Rosenbusch, Hannes
Published: (2025)
Overview of MWE history, challenges, and horizons: standing at the 20th anniversary of the MWE workshop series via MWE-UD2024
by: Han, Lifeng, et al.
Published: (2024)
by: Han, Lifeng, et al.
Published: (2024)
UD-KSL Treebank v1.3: A semi-automated framework for aligning XPOS-extracted units with UPOS tags
by: Sung, Hakyung, et al.
Published: (2025)
by: Sung, Hakyung, et al.
Published: (2025)
Does language matter for spoken word classification? A multilingual generative meta-learning approach
by: Ziki, Batsirayi Mupamhi, et al.
Published: (2026)
by: Ziki, Batsirayi Mupamhi, et al.
Published: (2026)
LLMs for automatic annotation of Mandarin narrative transcripts
by: Zhao, Qingwen, et al.
Published: (2026)
by: Zhao, Qingwen, et al.
Published: (2026)
Recovering document annotations for sentence-level bitext
by: Wicks, Rachel, et al.
Published: (2024)
by: Wicks, Rachel, et al.
Published: (2024)
Similar Items
-
The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices
by: Pannitto, Ludovica, et al.
Published: (2024) -
Towards the first UD Treebank of Spoken Italian: the KIParla forest
by: Pannitto, Ludovica
Published: (2024) -
Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
by: Simonotti, Martina, et al.
Published: (2026) -
Counting trees: A treebank-driven exploration of syntactic variation in speech and writing across languages
by: Dobrovoljc, Kaja
Published: (2025) -
Annotating Constructions with UD: the experience of the Italian Constructicon
by: Pannitto, Ludovica, et al.
Published: (2024)