FuLG: 150B Romanian Corpus for Language Model Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bădoiu, Vlad-Andrei, Dumitru, Mihai-Valentin, Gherghescu, Alexandru M., Agache, Alexandru, Raiciu, Costin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLMic: Romanian Foundation Language Model
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2025)
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2025)
I've Got 99 Problems But FLOPS Ain't One
von: Gherghescu, Alexandru M., et al.
Veröffentlicht: (2024)
von: Gherghescu, Alexandru M., et al.
Veröffentlicht: (2024)
Prose-to-P4: Leveraging High Level Languages
von: Dumitru, Mihai-Valentin, et al.
Veröffentlicht: (2024)
von: Dumitru, Mihai-Valentin, et al.
Veröffentlicht: (2024)
Enhancing Romanian Offensive Language Detection through Knowledge Distillation, Multi-Task Learning, and Data Augmentation
von: Matei, Vlad-Cristian, et al.
Veröffentlicht: (2024)
von: Matei, Vlad-Cristian, et al.
Veröffentlicht: (2024)
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
von: Man, Andrei Vlad, et al.
Veröffentlicht: (2025)
von: Man, Andrei Vlad, et al.
Veröffentlicht: (2025)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
von: Negoita, Vlad, et al.
Veröffentlicht: (2025)
von: Negoita, Vlad, et al.
Veröffentlicht: (2025)
RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2026)
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2026)
MuSaRoNews: A Multidomain, Multimodal Satire Dataset from Romanian News Articles
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
SaRoHead: Detecting Satire in a Multi-Domain Romanian News Headline Dataset
von: Vîrlan, Mihnea-Alexandru, et al.
Veröffentlicht: (2025)
von: Vîrlan, Mihnea-Alexandru, et al.
Veröffentlicht: (2025)
RoCoISLR: A Romanian Corpus for Isolated Sign Language Recognition
von: Rîpanu, Cătălin-Alexandru, et al.
Veröffentlicht: (2025)
von: Rîpanu, Cătălin-Alexandru, et al.
Veröffentlicht: (2025)
HistNERo: Historical Named Entity Recognition for the Romanian Language
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2024)
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2024)
RoQLlama: A Lightweight Romanian Adapted Language Model
von: Dima, George-Andrei, et al.
Veröffentlicht: (2024)
von: Dima, George-Andrei, et al.
Veröffentlicht: (2024)
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering
von: Crăciun, Cristian-George, et al.
Veröffentlicht: (2024)
von: Crăciun, Cristian-George, et al.
Veröffentlicht: (2024)
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
von: Dima, George-Andrei, et al.
Veröffentlicht: (2025)
von: Dima, George-Andrei, et al.
Veröffentlicht: (2025)
SeLeRoSa: Sentence-Level Romanian Satire Detection Dataset
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2025)
Reddit is all you need: Authorship profiling for Romanian
von: Ştefănescu, Ecaterina, et al.
Veröffentlicht: (2024)
von: Ştefănescu, Ecaterina, et al.
Veröffentlicht: (2024)
Multilingual Vision-Language Models, A Survey
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
von: Dumitru, Alexandru, et al.
Veröffentlicht: (2025)
von: Dumitru, Alexandru, et al.
Veröffentlicht: (2025)
Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
von: Manea, Andrei-Alexandru, et al.
Veröffentlicht: (2025)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
von: Masala, Mihai, et al.
Veröffentlicht: (2026)
von: Masala, Mihai, et al.
Veröffentlicht: (2026)
"Vorbeşti Româneşte?" A Recipe to Train Powerful Romanian LLMs with English Instructions
von: Masala, Mihai, et al.
Veröffentlicht: (2024)
von: Masala, Mihai, et al.
Veröffentlicht: (2024)
Querying Structured Data Through Natural Language Using Language Models
von: Valentin-Micu, Hontan, et al.
Veröffentlicht: (2026)
von: Valentin-Micu, Hontan, et al.
Veröffentlicht: (2026)
TF3-RO-50M: Training Compact Romanian Language Models from Scratch on Synthetic Moral Microfiction
von: Nadas, Mihai Dan, et al.
Veröffentlicht: (2026)
von: Nadas, Mihai Dan, et al.
Veröffentlicht: (2026)
RoMemes: A multimodal meme corpus for the Romanian language
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
MoRoVoc: A Large Dataset for Geographical Variation Identification of the Spoken Romanian Language
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2025)
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2025)
Evaluating Large Language Models for Diacritic Restoration in Romanian Texts: A Comparative Study
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
Investigating the Impact of Semi-Supervised Methods with Data Augmentation on Offensive Language Detection in Romanian Language
von: Nicola, Elena-Beatrice, et al.
Veröffentlicht: (2024)
von: Nicola, Elena-Beatrice, et al.
Veröffentlicht: (2024)
Investigating Large Language Models for Complex Word Identification in Multilingual and Multidomain Setups
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2024)
von: Smădu, Răzvan-Alexandru, et al.
Veröffentlicht: (2024)
A Culturally-Rich Romanian NLP Dataset from "Who Wants to Be a Millionaire?" Videos
von: Ganea, Alexandru-Gabriel, et al.
Veröffentlicht: (2025)
von: Ganea, Alexandru-Gabriel, et al.
Veröffentlicht: (2025)
RoLargeSum: A Large Dialect-Aware Romanian News Dataset for Summary, Headline, and Keyword Generation
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2024)
von: Avram, Andrei-Marius, et al.
Veröffentlicht: (2024)
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
von: Nadas, Mihai, et al.
Veröffentlicht: (2025)
From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples
von: Vacareanu, Robert, et al.
Veröffentlicht: (2024)
von: Vacareanu, Robert, et al.
Veröffentlicht: (2024)
Neural Grammatical Error Correction for Romanian
von: Cotet, Teodor-Mihai, et al.
Veröffentlicht: (2026)
von: Cotet, Teodor-Mihai, et al.
Veröffentlicht: (2026)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
von: Phatthiyaphaibun, Wannaphong, et al.
Veröffentlicht: (2025)
von: Phatthiyaphaibun, Wannaphong, et al.
Veröffentlicht: (2025)
RoBiologyDataChoiceQA: A Romanian Dataset for improving Biology understanding of Large Language Models
von: Ghinea, Dragos-Dumitru, et al.
Veröffentlicht: (2025)
von: Ghinea, Dragos-Dumitru, et al.
Veröffentlicht: (2025)
PsihoRo: Depression and Anxiety Romanian Text Corpus
von: Ciobotaru, Alexandra, et al.
Veröffentlicht: (2026)
von: Ciobotaru, Alexandra, et al.
Veröffentlicht: (2026)
Exploring the Impact of Corpus Diversity on Financial Pretrained Language Models
von: Choe, Jaeyoung, et al.
Veröffentlicht: (2023)
von: Choe, Jaeyoung, et al.
Veröffentlicht: (2023)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
von: Soldaini, Luca, et al.
Veröffentlicht: (2024)
von: Soldaini, Luca, et al.
Veröffentlicht: (2024)
Static Analysis Framework for Detecting Use-After-Free Bugs in C++
von: Teodorescu, Vlad-Alexandru, et al.
Veröffentlicht: (2024)
von: Teodorescu, Vlad-Alexandru, et al.
Veröffentlicht: (2024)
RELATE: A Modern Processing Platform for Romanian Language
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
von: Păiş, Vasile, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLMic: Romanian Foundation Language Model
von: Bădoiu, Vlad-Andrei, et al.
Veröffentlicht: (2025) -
I've Got 99 Problems But FLOPS Ain't One
von: Gherghescu, Alexandru M., et al.
Veröffentlicht: (2024) -
Prose-to-P4: Leveraging High Level Languages
von: Dumitru, Mihai-Valentin, et al.
Veröffentlicht: (2024) -
Enhancing Romanian Offensive Language Detection through Knowledge Distillation, Multi-Task Learning, and Data Augmentation
von: Matei, Vlad-Cristian, et al.
Veröffentlicht: (2024) -
RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams
von: Man, Andrei Vlad, et al.
Veröffentlicht: (2025)