Wiki Dumps to Training Corpora: South Slavic Case
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Škorić, Mihailo, Palma, Cosimo |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
New Textual Corpora for Serbian Language Modeling
par: Škorić, Mihailo, et autres
Publié: (2024)
par: Škorić, Mihailo, et autres
Publié: (2024)
Novi jezički modeli za srpski jezik
par: Škorić, Mihailo
Publié: (2024)
par: Škorić, Mihailo
Publié: (2024)
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026)
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024)
par: Ljubešić, Nikola, et autres
Publié: (2024)
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
par: Göpfert, Jan, et autres
Publié: (2025)
par: Göpfert, Jan, et autres
Publié: (2025)
API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMs
par: Basu, Kinjal, et autres
Publié: (2024)
par: Basu, Kinjal, et autres
Publié: (2024)
EvoWiki: Evaluating LLMs on Evolving Knowledge
par: Tang, Wei, et autres
Publié: (2024)
par: Tang, Wei, et autres
Publié: (2024)
WiCER: Wiki-memory Compile, Evaluate, Refine Iterative Knowledge Compilation for LLM Wiki Systems
par: Huerta, Juan M.
Publié: (2026)
par: Huerta, Juan M.
Publié: (2026)
WikiSplit++: Easy Data Refinement for Split and Rephrase
par: Tsukagoshi, Hayato, et autres
Publié: (2024)
par: Tsukagoshi, Hayato, et autres
Publié: (2024)
Validating and Exploring Large Geographic Corpora
par: Dunn, Jonathan
Publié: (2024)
par: Dunn, Jonathan
Publié: (2024)
A Review of the Challenges with Massive Web-mined Corpora Used in Large Language Models Pre-Training
par: Perełkiewicz, Michał, et autres
Publié: (2024)
par: Perełkiewicz, Michał, et autres
Publié: (2024)
Identifying Emerging Concepts in Large Corpora
par: Ma, Sibo, et autres
Publié: (2025)
par: Ma, Sibo, et autres
Publié: (2025)
PumpSense: Real-Time Detection and Target Extraction of Crypto Pump-and-Dumps on Telegram
par: Mahrous, Ahmed, et autres
Publié: (2026)
par: Mahrous, Ahmed, et autres
Publié: (2026)
Cross-lingual Named Entity Corpus for Slavic Languages
par: Piskorski, Jakub, et autres
Publié: (2024)
par: Piskorski, Jakub, et autres
Publié: (2024)
A multilingual hallucination benchmark: MultiWikiQHalluA
par: Thoresen, Freja, et autres
Publié: (2026)
par: Thoresen, Freja, et autres
Publié: (2026)
Bias in News Summarization: Measures, Pitfalls and Corpora
par: Steen, Julius, et autres
Publié: (2023)
par: Steen, Julius, et autres
Publié: (2023)
Comparable Corpora: Opportunities for New Research Directions
par: Church, Kenneth
Publié: (2025)
par: Church, Kenneth
Publié: (2025)
Transfer Learning for an Endangered Slavic Variety: Dependency Parsing in Pomak Across Contact-Shaped Dialects
par: Karakaş, Sercan
Publié: (2026)
par: Karakaş, Sercan
Publié: (2026)
Mathematical Entities: Corpora and Benchmarks
par: Collard, Jacob, et autres
Publié: (2024)
par: Collard, Jacob, et autres
Publié: (2024)
MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languages
par: Smart, Dan Saattrup
Publié: (2025)
par: Smart, Dan Saattrup
Publié: (2025)
AI Brown and AI Koditex: LLM-Generated Corpora Comparable to Traditional Corpora of English and Czech Texts
par: Milička, Jiří, et autres
Publié: (2025)
par: Milička, Jiří, et autres
Publié: (2025)
WikiVideo: Article Generation from Multiple Videos
par: Martin, Alexander, et autres
Publié: (2025)
par: Martin, Alexander, et autres
Publié: (2025)
Disambiguating Numeral Sequences to Decipher Ancient Accounting Corpora
par: Born, Logan, et autres
Publié: (2025)
par: Born, Logan, et autres
Publié: (2025)
Active Learning for Multilingual Fingerspelling Corpora
par: Wang, Shuai, et autres
Publié: (2023)
par: Wang, Shuai, et autres
Publié: (2023)
Building and Aligning Comparable Corpora
par: Saad, Motaz, et autres
Publié: (2025)
par: Saad, Motaz, et autres
Publié: (2025)
Guylingo: The Republic of Guyana Creole Corpora
par: Clarke, Christopher, et autres
Publié: (2024)
par: Clarke, Christopher, et autres
Publié: (2024)
Unsupervised Location Mapping for Narrative Corpora
par: Wagner, Eitan, et autres
Publié: (2025)
par: Wagner, Eitan, et autres
Publié: (2025)
Preference Consistency Matters: Enhancing Preference Learning in Language Models with Automated Self-Curation of Training Corpora
par: Lee, JoonHo, et autres
Publié: (2024)
par: Lee, JoonHo, et autres
Publié: (2024)
Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki
par: Ming, Haoliang, et autres
Publié: (2026)
par: Ming, Haoliang, et autres
Publié: (2026)
WikiSQE: A Large-Scale Dataset for Sentence Quality Estimation in Wikipedia
par: Ando, Kenichiro, et autres
Publié: (2023)
par: Ando, Kenichiro, et autres
Publié: (2023)
JGU Mainz's Submission to the WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA
par: Saadi, Hossain Shaikh, et autres
Publié: (2025)
par: Saadi, Hossain Shaikh, et autres
Publié: (2025)
Multilingual Embedding Probes Fail to Generalize Across Learner Corpora
par: Lyngbaek, Laurits, et autres
Publié: (2026)
par: Lyngbaek, Laurits, et autres
Publié: (2026)
HistLens: Mapping Idea Change across Concepts and Corpora
par: Jing, Yi, et autres
Publié: (2026)
par: Jing, Yi, et autres
Publié: (2026)
Data Caricatures: On the Representation of African American Language in Pretraining Corpora
par: Deas, Nicholas, et autres
Publié: (2025)
par: Deas, Nicholas, et autres
Publié: (2025)
AcrosticSleuth: Probabilistic Identification and Ranking of Acrostics in Multilingual Corpora
par: Fedchin, Aleksandr, et autres
Publié: (2024)
par: Fedchin, Aleksandr, et autres
Publié: (2024)
MASRAD: Arabic Terminology Management Corpora with Semi-Automatic Construction
par: Nasser, Mahdi, et autres
Publié: (2025)
par: Nasser, Mahdi, et autres
Publié: (2025)
When Should Dense Retrievers Be Updated in Evolving Corpora? Detecting Out-of-Distribution Corpora Using GradNormIR
par: Ko, Dayoon, et autres
Publié: (2025)
par: Ko, Dayoon, et autres
Publié: (2025)
Entity Insertion in Multilingual Linked Corpora: The Case of Wikipedia
par: Feith, Tomás, et autres
Publié: (2024)
par: Feith, Tomás, et autres
Publié: (2024)
LLMSQL: Upgrading WikiSQL for the LLM Era of Text-to-SQL
par: Pihulski, Dzmitry, et autres
Publié: (2025)
par: Pihulski, Dzmitry, et autres
Publié: (2025)
Documents similaires
-
New Textual Corpora for Serbian Language Modeling
par: Škorić, Mihailo, et autres
Publié: (2024) -
Novi jezički modeli za srpski jezik
par: Škorić, Mihailo
Publié: (2024) -
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
par: Pungeršek, Taja Kuzman, et autres
Publié: (2026) -
CLASSLA-web: Comparable Web Corpora of South Slavic Languages Enriched with Linguistic and Genre Annotation
par: Ljubešić, Nikola, et autres
Publié: (2024) -
State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?
par: Pungeršek, Taja Kuzman, et autres
Publié: (2025)