Is Semi-Automatic Transcription Useful in Corpus Creation? Preliminary Considerations on the KIParla Corpus
Fuente:
arXiv
Saved in:
| Main Authors: | Simonotti, Martina, Pannitto, Ludovica, Zucchini, Eleonora, Ballarè, Silvia, Mauri, Caterina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards the first UD Treebank of Spoken Italian: the KIParla forest
by: Pannitto, Ludovica
Published: (2024)
by: Pannitto, Ludovica
Published: (2024)
The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Coconstructions in spoken data: UD annotation guidelines and first results
by: Pannitto, Ludovica, et al.
Published: (2026)
by: Pannitto, Ludovica, et al.
Published: (2026)
Recurrent babbling: evaluating the acquisition of grammar from limited input data
by: Pannitto, Ludovica, et al.
Published: (2020)
by: Pannitto, Ludovica, et al.
Published: (2020)
CALaMo: a Constructionist Assessment of Language Models
by: Pannitto, Ludovica, et al.
Published: (2023)
by: Pannitto, Ludovica, et al.
Published: (2023)
'Layer su Layer': Identifying and Disambiguating the Italian NPN Construction in BERT's family
by: Gorzoni, Greta, et al.
Published: (2026)
by: Gorzoni, Greta, et al.
Published: (2026)
Did somebody say "Gest-IT"? A pilot exploration of multimodal data management
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Corpus Considerations for Annotator Modeling and Scaling
by: Sarumi, Olufunke O., et al.
Published: (2024)
by: Sarumi, Olufunke O., et al.
Published: (2024)
Constraining constructions with WordNet: pros and cons for the semantic annotation of fillers in the Italian Constructicon
by: Pisciotta, Flavio, et al.
Published: (2025)
by: Pisciotta, Flavio, et al.
Published: (2025)
English to Central Kurdish Speech Translation: Corpus Creation, Evaluation, and Orthographic Standardization
by: Mohammadamini, Mohammad, et al.
Published: (2026)
by: Mohammadamini, Mohammad, et al.
Published: (2026)
DEBISS: a Corpus of Individual, Semi-structured and Spoken Debates
by: de Souza, Klaywert Danillo Ferreira, et al.
Published: (2026)
by: de Souza, Klaywert Danillo Ferreira, et al.
Published: (2026)
Annotating Constructions with UD: the experience of the Italian Constructicon
by: Pannitto, Ludovica, et al.
Published: (2024)
by: Pannitto, Ludovica, et al.
Published: (2024)
Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
by: D, Anusha M, et al.
Published: (2025)
by: D, Anusha M, et al.
Published: (2025)
Creating an Aligned Corpus of Sound and Text: The Multimodal Corpus of Shakespeare and Milton
by: Agirrezabal, Manex
Published: (2024)
by: Agirrezabal, Manex
Published: (2024)
Automatic Construction of a Large-Scale Corpus for Geoparsing Using Wikipedia Hyperlinks
by: Ohno, Keyaki, et al.
Published: (2024)
by: Ohno, Keyaki, et al.
Published: (2024)
The Material Contracts Corpus
by: Adelson, Peter, et al.
Published: (2025)
by: Adelson, Peter, et al.
Published: (2025)
The Russian Legislative Corpus
by: Saveliev, Denis, et al.
Published: (2024)
by: Saveliev, Denis, et al.
Published: (2024)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
Distinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
by: Arpa, Zaara Zabeen, et al.
Published: (2025)
by: Arpa, Zaara Zabeen, et al.
Published: (2025)
Swiss Parliaments Corpus Re-Imagined (SPC_R): Enhanced Transcription with RAG-based Correction and Predicted BLEU
by: Timmel, Vincenzo, et al.
Published: (2025)
by: Timmel, Vincenzo, et al.
Published: (2025)
Investigating Affect Mining Techniques for Annotation Sample Selection in the Creation of Finnish Affective Speech Corpus
by: Lahtinen, Kalle, et al.
Published: (2025)
by: Lahtinen, Kalle, et al.
Published: (2025)
The Medical Metaphors Corpus (MCC)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
by: Lippolis, Anna Sofia, et al.
Published: (2025)
Speak & Improve Corpus 2025: an L2 English Speech Corpus for Language Assessment and Feedback
by: Knill, Kate, et al.
Published: (2024)
by: Knill, Kate, et al.
Published: (2024)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
by: Goldin, Gili, et al.
Published: (2024)
by: Goldin, Gili, et al.
Published: (2024)
NLAS-multi: A Multilingual Corpus of Automatically Generated Natural Language Argumentation Schemes
by: Ruiz-Dolz, Ramon, et al.
Published: (2024)
by: Ruiz-Dolz, Ramon, et al.
Published: (2024)
CorpusQA: A 10 Million Token Benchmark for Corpus-Level Analysis and Reasoning
by: Lu, Zhiyuan, et al.
Published: (2026)
by: Lu, Zhiyuan, et al.
Published: (2026)
CorpusLM: Towards a Unified Language Model on Corpus for Knowledge-Intensive Tasks
by: Li, Xiaoxi, et al.
Published: (2024)
by: Li, Xiaoxi, et al.
Published: (2024)
The TUB Sign Language Corpus Collection
by: Avramidis, Eleftherios, et al.
Published: (2025)
by: Avramidis, Eleftherios, et al.
Published: (2025)
The Pilot Corpus of the English Semantic Sketches
by: Petrova, Maria, et al.
Published: (2025)
by: Petrova, Maria, et al.
Published: (2025)
The SAMER Arabic Text Simplification Corpus
by: Alhafni, Bashar, et al.
Published: (2024)
by: Alhafni, Bashar, et al.
Published: (2024)
The Moral Foundations Reddit Corpus
by: Trager, Jackson, et al.
Published: (2022)
by: Trager, Jackson, et al.
Published: (2022)
The PLLuM Instruction Corpus
by: Pęzik, Piotr, et al.
Published: (2025)
by: Pęzik, Piotr, et al.
Published: (2025)
The Moral Foundations Weibo Corpus
by: Cao, Renjie, et al.
Published: (2024)
by: Cao, Renjie, et al.
Published: (2024)
The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
by: Lequeu, Pierre-Antoine, et al.
Published: (2026)
Speech Corpus for Korean Children with Autism Spectrum Disorder: Towards Automatic Assessment Systems
by: Lee, Seonwoo, et al.
Published: (2024)
by: Lee, Seonwoo, et al.
Published: (2024)
Extending a Parliamentary Corpus with MPs' Tweets: Automatic Annotation and Evaluation Using MultiParTweet
by: Bagci, Mevlüt, et al.
Published: (2025)
by: Bagci, Mevlüt, et al.
Published: (2025)
A French Version of the OLDI Seed Corpus
by: Marmonier, Malik, et al.
Published: (2025)
by: Marmonier, Malik, et al.
Published: (2025)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
SiDiaC: Sinhala Diachronic Corpus
by: Jayatilleke, Nevidu, et al.
Published: (2025)
by: Jayatilleke, Nevidu, et al.
Published: (2025)
FFSTC: Fongbe to French Speech Translation Corpus
by: Kponou, D. Fortune, et al.
Published: (2024)
by: Kponou, D. Fortune, et al.
Published: (2024)
Similar Items
-
Towards the first UD Treebank of Spoken Italian: the KIParla forest
by: Pannitto, Ludovica
Published: (2024) -
The KIPARLA Forest treebank of spoken Italian: an overview of initial design choices
by: Pannitto, Ludovica, et al.
Published: (2024) -
Coconstructions in spoken data: UD annotation guidelines and first results
by: Pannitto, Ludovica, et al.
Published: (2026) -
Recurrent babbling: evaluating the acquisition of grammar from limited input data
by: Pannitto, Ludovica, et al.
Published: (2020) -
CALaMo: a Constructionist Assessment of Language Models
by: Pannitto, Ludovica, et al.
Published: (2023)