Neural spell-checker: Beyond words with synthetic data generation
Fuente:
arXiv
Saved in:
| Main Authors: | Klemen, Matej, Božič, Martin, Holdt, Špela Arhar, Robnik-Šikonja, Marko |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
by: Arčon, Tjaša, et al.
Published: (2026)
by: Arčon, Tjaša, et al.
Published: (2026)
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026)
by: Miok, Kristian, et al.
Published: (2026)
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
by: Klemen, Matej, et al.
Published: (2025)
by: Klemen, Matej, et al.
Published: (2025)
Code-mixed Sentiment and Hate-speech Prediction
by: Yadav, Anjali, et al.
Published: (2024)
by: Yadav, Anjali, et al.
Published: (2024)
Generative Model for Less-Resourced Language with 1 billion parameters
by: Vreš, Domen, et al.
Published: (2024)
by: Vreš, Domen, et al.
Published: (2024)
Solving Word-Sense Disambiguation and Word-Sense Induction with Dictionary Examples
by: Škvorc, Tadej, et al.
Published: (2025)
by: Škvorc, Tadej, et al.
Published: (2025)
QFS-Composer: Query-focused summarization pipeline for less resourced languages
by: Đuranović, Vuk, et al.
Published: (2026)
by: Đuranović, Vuk, et al.
Published: (2026)
Sarcasm Detection in a Less-Resourced Language
by: Đoković, Lazar, et al.
Published: (2024)
by: Đoković, Lazar, et al.
Published: (2024)
Improving LLMs for Machine Translation Using Synthetic Preference Data
by: Vajda, Dario, et al.
Published: (2025)
by: Vajda, Dario, et al.
Published: (2025)
Large language models for folktale type automation based on motifs: Cinderella case study
by: Arčon, Tjaša, et al.
Published: (2025)
by: Arčon, Tjaša, et al.
Published: (2025)
Real-time News Story Identification
by: Škvorc, Tadej, et al.
Published: (2025)
by: Škvorc, Tadej, et al.
Published: (2025)
Review of Natural Language Processing in Pharmacology
by: Trajanov, Dimitar, et al.
Published: (2022)
by: Trajanov, Dimitar, et al.
Published: (2022)
Measuring Catastrophic Forgetting in Cross-Lingual Transfer Paradigms: Exploring Tuning Strategies
by: Koloski, Boshko, et al.
Published: (2023)
by: Koloski, Boshko, et al.
Published: (2023)
TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
by: Miok, Kristian, et al.
Published: (2025)
by: Miok, Kristian, et al.
Published: (2025)
Incremental Graph Construction Enables Robust Spectral Clustering of Texts
by: Pranjić, Marko, et al.
Published: (2026)
by: Pranjić, Marko, et al.
Published: (2026)
The truth is no diaper: Human and AI-generated associations to emotional words
by: Vintar, Špela, et al.
Published: (2025)
by: Vintar, Špela, et al.
Published: (2025)
Retrieval-augmented code completion for local projects using large language models
by: Hostnik, Marko, et al.
Published: (2024)
by: Hostnik, Marko, et al.
Published: (2024)
Building a Strong Instruction Language Model for a Less-Resourced Language
by: Vreš, Domen, et al.
Published: (2026)
by: Vreš, Domen, et al.
Published: (2026)
Semantics or spelling? Probing contextual word embeddings with orthographic noise
by: Matthews, Jacob A., et al.
Published: (2024)
by: Matthews, Jacob A., et al.
Published: (2024)
A symbolic Perl algorithm for the unification of Nahuatl word spellings
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
by: Guzmán-Landa, Juan-José, et al.
Published: (2025)
$statcheck$ is flawed by design and no valid spell checker for statistical results
by: Böschen, Ingmar
Published: (2024)
by: Böschen, Ingmar
Published: (2024)
A Survey of Deep Learning Audio Generation Methods
by: Božić, Matej, et al.
Published: (2024)
by: Božić, Matej, et al.
Published: (2024)
CLIPPER: Compression enables long-context synthetic data generation
by: Pham, Chau Minh, et al.
Published: (2025)
by: Pham, Chau Minh, et al.
Published: (2025)
Factcheck-Bench: Fine-Grained Evaluation Benchmark for Automatic Fact-checkers
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
From Polyester Girlfriends to Blind Mice: Creating the First Pragmatics Understanding Benchmarks for Slovene
by: Brglez, Mojca, et al.
Published: (2025)
by: Brglez, Mojca, et al.
Published: (2025)
Provenance: A Light-weight Fact-checker for Retrieval Augmented LLM Generation Output
by: Sankararaman, Hithesh, et al.
Published: (2024)
by: Sankararaman, Hithesh, et al.
Published: (2024)
Zero-shot generation of synthetic neurosurgical data with large language models
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
LLMs' morphological analyses of complex FST-generated Finnish words
by: Moisio, Anssi, et al.
Published: (2024)
by: Moisio, Anssi, et al.
Published: (2024)
WebJspell an online morphological analyser and spell checker
by: Rui Vilela
Published: (2007)
by: Rui Vilela
Published: (2007)
Human-centered NLP Fact-checking: Co-Designing with Fact-checkers using Matchmaking for AI
by: Liu, Houjiang, et al.
Published: (2023)
by: Liu, Houjiang, et al.
Published: (2023)
Tracking Semantic Change in Slovene: A Novel Dataset and Optimal Transport-Based Distance
by: Pranjić, Marko, et al.
Published: (2024)
by: Pranjić, Marko, et al.
Published: (2024)
GPT Assisted Annotation of Rhetorical and Linguistic Features for Interpretable Propaganda Technique Detection in News Text
by: Hamilton, Kyle, et al.
Published: (2024)
by: Hamilton, Kyle, et al.
Published: (2024)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
synthocr-gen: A synthetic ocr dataset generator for low-resource languages- breaking the data barrier
by: Malik, Haq Nawaz, et al.
Published: (2026)
by: Malik, Haq Nawaz, et al.
Published: (2026)
LLMCARE: early detection of cognitive impairment via transformer models enhanced by LLM-generated synthetic data
by: Zolnour, Ali, et al.
Published: (2025)
by: Zolnour, Ali, et al.
Published: (2025)
Optimal word order for non-causal text generation with Large Language Models: the Spanish case
by: Busto-Castiñeira, Andrea, et al.
Published: (2025)
by: Busto-Castiñeira, Andrea, et al.
Published: (2025)
Cover Image, Volume 121, Number 5, May 2024
by: Klemen Božič, et al.
Published: (2024)
by: Klemen Božič, et al.
Published: (2024)
Scaling few-shot spoken word classification with generative meta-continual learning
by: Beyers, Louise, et al.
Published: (2026)
by: Beyers, Louise, et al.
Published: (2026)
Generative adversarial networks vs large language models: a comparative study on synthetic tabular data generation
by: Barr, Austin A., et al.
Published: (2025)
by: Barr, Austin A., et al.
Published: (2025)
Representing data in words: A context engineering approach
by: Caut, Amandine M., et al.
Published: (2025)
by: Caut, Amandine M., et al.
Published: (2025)
Similar Items
-
Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages
by: Arčon, Tjaša, et al.
Published: (2026) -
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026) -
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
by: Klemen, Matej, et al.
Published: (2025) -
Code-mixed Sentiment and Hate-speech Prediction
by: Yadav, Anjali, et al.
Published: (2024) -
Generative Model for Less-Resourced Language with 1 billion parameters
by: Vreš, Domen, et al.
Published: (2024)