Autocorrect for Estonian texts: final report from project EKTB25
Fuente:
arXiv
Guardado en:
| Autores principales: | Luhtaru, Agnes, Vainikko, Martin, Liin, Krista, Allkivi-Metsoja, Kais, Kippar, Jaagup, Eslon, Pille, Fishel, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
por: Allkivi, Kais
Publicado: (2026)
por: Allkivi, Kais
Publicado: (2026)
To Err Is Human, but Llamas Can Learn It Too
por: Luhtaru, Agnes, et al.
Publicado: (2024)
por: Luhtaru, Agnes, et al.
Publicado: (2024)
Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer
por: Kuulmets, Hele-Andra, et al.
Publicado: (2024)
por: Kuulmets, Hele-Andra, et al.
Publicado: (2024)
Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
por: Karjus, Andres, et al.
Publicado: (2026)
por: Karjus, Andres, et al.
Publicado: (2026)
Limited Linguistic Diversity in Embodied AI Datasets
por: Wanna, Selma, et al.
Publicado: (2026)
por: Wanna, Selma, et al.
Publicado: (2026)
EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Training
por: Dorkin, Aleksei, et al.
Publicado: (2026)
por: Dorkin, Aleksei, et al.
Publicado: (2026)
Optimizing Estonian TV Subtitles with Semi-supervised Learning and LLMs
por: Fedorchenko, Artem, et al.
Publicado: (2025)
por: Fedorchenko, Artem, et al.
Publicado: (2025)
How Uncertainty Estimation Scales with Sampling in Reasoning Models
por: Del, Maksym, et al.
Publicado: (2026)
por: Del, Maksym, et al.
Publicado: (2026)
Are generative AI text annotations systematically biased?
por: Stolwijk, Sjoerd B., et al.
Publicado: (2025)
por: Stolwijk, Sjoerd B., et al.
Publicado: (2025)
LLMs for Extremely Low-Resource Finno-Ugric Languages
por: Purason, Taido, et al.
Publicado: (2024)
por: Purason, Taido, et al.
Publicado: (2024)
Estonian Native Large Language Model Benchmark
por: Lillepalu, Helena Grete, et al.
Publicado: (2025)
por: Lillepalu, Helena Grete, et al.
Publicado: (2025)
Prune or Retrain: Optimizing the Vocabulary of Multilingual Models for Estonian
por: Dorkin, Aleksei, et al.
Publicado: (2025)
por: Dorkin, Aleksei, et al.
Publicado: (2025)
CAPRAG: A Large Language Model Solution for Customer Service and Automatic Reporting using Vector and Graph Retrieval-Augmented Generation
por: Landolsi, Hamza, et al.
Publicado: (2025)
por: Landolsi, Hamza, et al.
Publicado: (2025)
Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
por: Purason, Taido, et al.
Publicado: (2025)
por: Purason, Taido, et al.
Publicado: (2025)
Abusive text transformation using LLMs
por: Chandra, Rohitash, et al.
Publicado: (2025)
por: Chandra, Rohitash, et al.
Publicado: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
por: Miralles-González, Pablo, et al.
Publicado: (2025)
por: Miralles-González, Pablo, et al.
Publicado: (2025)
Benchmark of stylistic variation in LLM-generated texts
por: Milička, Jiří, et al.
Publicado: (2025)
por: Milička, Jiří, et al.
Publicado: (2025)
PAGE: Prompt Augmentation for text Generation Enhancement
por: Pacchiotti, Mauro Jose, et al.
Publicado: (2025)
por: Pacchiotti, Mauro Jose, et al.
Publicado: (2025)
Serialized EHR make for good text representations
por: Chou, Zhirong, et al.
Publicado: (2025)
por: Chou, Zhirong, et al.
Publicado: (2025)
An experimental and computational study of an Estonian single-person word naming
por: Lõo, Kaidi, et al.
Publicado: (2025)
por: Lõo, Kaidi, et al.
Publicado: (2025)
Comparison of Current Approaches to Lemmatization: A Case Study in Estonian
por: Dorkin, Aleksei, et al.
Publicado: (2024)
por: Dorkin, Aleksei, et al.
Publicado: (2024)
GliLem: Leveraging GliNER for Contextualized Lemmatization in Estonian
por: Dorkin, Aleksei, et al.
Publicado: (2024)
por: Dorkin, Aleksei, et al.
Publicado: (2024)
From Reasoning to Generalization: Knowledge-Augmented LLMs for ARC Benchmark
por: Lei, Chao, et al.
Publicado: (2025)
por: Lei, Chao, et al.
Publicado: (2025)
Large language models struggle with ethnographic text annotation
por: Goodall, Leonardo S., et al.
Publicado: (2026)
por: Goodall, Leonardo S., et al.
Publicado: (2026)
Beyond checkmate: exploring the creative chokepoints in AI text
por: Tripto, Nafis Irtiza, et al.
Publicado: (2025)
por: Tripto, Nafis Irtiza, et al.
Publicado: (2025)
Can professional translators identify machine-generated text?
por: Farrell, Michael
Publicado: (2026)
por: Farrell, Michael
Publicado: (2026)
LLMs can hide text in other text of the same length
por: Norelli, Antonio, et al.
Publicado: (2025)
por: Norelli, Antonio, et al.
Publicado: (2025)
A Survey on Image-text Multimodal Models
por: Guo, Ruifeng, et al.
Publicado: (2023)
por: Guo, Ruifeng, et al.
Publicado: (2023)
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
por: Qian, Shun, et al.
Publicado: (2026)
por: Qian, Shun, et al.
Publicado: (2026)
Can postgraduate translation students identify machine-generated text?
por: Farrell, Michael
Publicado: (2025)
por: Farrell, Michael
Publicado: (2025)
Differentially-private text generation degrades output language quality
por: Çano, Erion, et al.
Publicado: (2025)
por: Çano, Erion, et al.
Publicado: (2025)
Detecting value-expressive text posts in Russian social media
por: Milkova, Maria, et al.
Publicado: (2023)
por: Milkova, Maria, et al.
Publicado: (2023)
Meta-aware Learning in text-to-SQL Large Language Model
por: Zhang, Wenda
Publicado: (2025)
por: Zhang, Wenda
Publicado: (2025)
Creation of the Estonian Subjectivity Dataset: Assessing the Degree of Subjectivity on a Scale
por: Gailit, Karl Gustav, et al.
Publicado: (2025)
por: Gailit, Karl Gustav, et al.
Publicado: (2025)
Multilingual transformer and BERTopic for short text topic modeling: The case of Serbian
por: Medvecki, Darija, et al.
Publicado: (2024)
por: Medvecki, Darija, et al.
Publicado: (2024)
ARC-Encoder: learning compressed text representations for large language models
por: Pilchen, Hippolyte, et al.
Publicado: (2025)
por: Pilchen, Hippolyte, et al.
Publicado: (2025)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
por: Tran, Hieu, et al.
Publicado: (2024)
por: Tran, Hieu, et al.
Publicado: (2024)
EzSQL: An SQL intermediate representation for improving SQL-to-text Generation
por: Bhardwaj, Meher, et al.
Publicado: (2024)
por: Bhardwaj, Meher, et al.
Publicado: (2024)
Improving the quality of Persian clinical text with a novel spelling correction system
por: Dashti, Seyed Mohammad Sadegh, et al.
Publicado: (2024)
por: Dashti, Seyed Mohammad Sadegh, et al.
Publicado: (2024)
A Chat About Boring Problems: Studying GPT-based text normalization
por: Zhang, Yang, et al.
Publicado: (2023)
por: Zhang, Yang, et al.
Publicado: (2023)
Ejemplares similares
-
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
por: Allkivi, Kais
Publicado: (2026) -
To Err Is Human, but Llamas Can Learn It Too
por: Luhtaru, Agnes, et al.
Publicado: (2024) -
Teaching Llama a New Language Through Cross-Lingual Knowledge Transfer
por: Kuulmets, Hele-Andra, et al.
Publicado: (2024) -
Machine-Assisted Grading of Nationwide School-Leaving Essay Exams with LLMs and Statistical NLP
por: Karjus, Andres, et al.
Publicado: (2026) -
Limited Linguistic Diversity in Embodied AI Datasets
por: Wanna, Selma, et al.
Publicado: (2026)