Salvato in:
| Autori principali: | van Oort, Jesse, Brinkkemper, Frank, de Graaf, Erik, Vanroy, Bram, Lensink, Saskia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.00920 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fietje: An open, efficient LLM for Dutch
di: Vanroy, Bram
Pubblicazione: (2024)
di: Vanroy, Bram
Pubblicazione: (2024)
GEITje 7B Ultra: A Conversational Model for Dutch
di: Vanroy, Bram
Pubblicazione: (2024)
di: Vanroy, Bram
Pubblicazione: (2024)
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
di: Raes, Rik, et al.
Pubblicazione: (2024)
di: Raes, Rik, et al.
Pubblicazione: (2024)
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
di: Banar, Nikolay, et al.
Pubblicazione: (2025)
di: Banar, Nikolay, et al.
Pubblicazione: (2025)
Biomedical Entity Linking for Dutch: Fine-tuning a Self-alignment BERT Model on an Automatically Generated Wikipedia Corpus
di: Hartendorp, Fons, et al.
Pubblicazione: (2024)
di: Hartendorp, Fons, et al.
Pubblicazione: (2024)
BEIR-NL: Zero-shot Information Retrieval Benchmark for the Dutch Language
di: Banar, Nikolay, et al.
Pubblicazione: (2024)
di: Banar, Nikolay, et al.
Pubblicazione: (2024)
PhoGPT: Generative Pre-training for Vietnamese
di: Nguyen, Dat Quoc, et al.
Pubblicazione: (2023)
di: Nguyen, Dat Quoc, et al.
Pubblicazione: (2023)
Enhancing Summarization Performance through Transformer-Based Prompt Engineering in Automated Medical Reporting
di: van Zandvoort, Daphne, et al.
Pubblicazione: (2023)
di: van Zandvoort, Daphne, et al.
Pubblicazione: (2023)
Comparative Experimentation of Accuracy Metrics in Automated Medical Reporting: The Case of Otitis Consultations
di: Faber, Wouter, et al.
Pubblicazione: (2023)
di: Faber, Wouter, et al.
Pubblicazione: (2023)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
di: Shekoufandeh, Golshid, et al.
Pubblicazione: (2025)
RecGPT: Generative Pre-training for Text-based Recommendation
di: Ngo, Hoang, et al.
Pubblicazione: (2024)
di: Ngo, Hoang, et al.
Pubblicazione: (2024)
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
di: Langlais, Pierre-Carl, et al.
Pubblicazione: (2025)
di: Langlais, Pierre-Carl, et al.
Pubblicazione: (2025)
Summarizing long regulatory documents with a multi-step pipeline
di: Sie, Mika, et al.
Pubblicazione: (2024)
di: Sie, Mika, et al.
Pubblicazione: (2024)
Pre-training and Diagnosing Knowledge Base Completion Models
di: Kocijan, Vid, et al.
Pubblicazione: (2024)
di: Kocijan, Vid, et al.
Pubblicazione: (2024)
Hybrid-NL2SVA: Integrating RAG and Finetuning for LLM-based NL2SVA
di: Xiao, Weihua, et al.
Pubblicazione: (2025)
di: Xiao, Weihua, et al.
Pubblicazione: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
di: Yadav, Sumit, et al.
Pubblicazione: (2025)
Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity
di: Jewitt, James, et al.
Pubblicazione: (2026)
di: Jewitt, James, et al.
Pubblicazione: (2026)
TransGPT: Multi-modal Generative Pre-trained Transformer for Transportation
di: Wang, Peng, et al.
Pubblicazione: (2024)
di: Wang, Peng, et al.
Pubblicazione: (2024)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
di: Joshi, Raviraj, et al.
Pubblicazione: (2024)
di: Joshi, Raviraj, et al.
Pubblicazione: (2024)
Phonological Neighbourhood Density in Dutch Verbs: From Classification to Corpus Annotation
di: Zimianiti, Eleni, et al.
Pubblicazione: (2025)
di: Zimianiti, Eleni, et al.
Pubblicazione: (2025)
Scaling LLM Pre-training with Vocabulary Curriculum
di: Yu, Fangyuan
Pubblicazione: (2025)
di: Yu, Fangyuan
Pubblicazione: (2025)
LicenseGPT: A Fine-tuned Foundation Model for Publicly Available Dataset License Compliance
di: Tan, Jingwen, et al.
Pubblicazione: (2024)
di: Tan, Jingwen, et al.
Pubblicazione: (2024)
QuCo-RAG: Quantifying Uncertainty from the Pre-training Corpus for Dynamic Retrieval-Augmented Generation
di: Min, Dehai, et al.
Pubblicazione: (2025)
di: Min, Dehai, et al.
Pubblicazione: (2025)
GPTEval: A Survey on Assessments of ChatGPT and GPT-4
di: Mao, Rui, et al.
Pubblicazione: (2023)
di: Mao, Rui, et al.
Pubblicazione: (2023)
NL2SQLBench: A Modular Benchmarking Framework for LLM-Enabled NL2SQL Solutions
di: Hou, Shizheng, et al.
Pubblicazione: (2026)
di: Hou, Shizheng, et al.
Pubblicazione: (2026)
NITP: Next Implicit Token Prediction for LLM Pre-training
di: Zhang, Xiangdong, et al.
Pubblicazione: (2026)
di: Zhang, Xiangdong, et al.
Pubblicazione: (2026)
Dutch CrowS-Pairs: Adapting a Challenge Dataset for Measuring Social Biases in Language Models for Dutch
di: Strazda, Elza, et al.
Pubblicazione: (2025)
di: Strazda, Elza, et al.
Pubblicazione: (2025)
SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
di: Zhu, Rui-Jie, et al.
Pubblicazione: (2023)
di: Zhu, Rui-Jie, et al.
Pubblicazione: (2023)
Chapter Bridging scaling with agglomeration economies
di: van Oort, Frank G.
Pubblicazione: (2024)
di: van Oort, Frank G.
Pubblicazione: (2024)
More on Maximally Permissive Similarity Control of Discrete Event Systems
di: Wang, Yu, et al.
Pubblicazione: (2024)
di: Wang, Yu, et al.
Pubblicazione: (2024)
Evaluating LLMs and Pre-trained Models for Text Summarization Across Diverse Datasets
di: Rehman, Tohida, et al.
Pubblicazione: (2025)
di: Rehman, Tohida, et al.
Pubblicazione: (2025)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
di: Nguyen, Huu, et al.
Pubblicazione: (2025)
di: Nguyen, Huu, et al.
Pubblicazione: (2025)
Diagnosis extraction from unstructured Dutch echocardiogram reports using span- and document-level characteristic classification
di: Arends, Bauke, et al.
Pubblicazione: (2024)
di: Arends, Bauke, et al.
Pubblicazione: (2024)
SparseLLM: Towards Global Pruning for Pre-trained Language Models
di: Bai, Guangji, et al.
Pubblicazione: (2024)
di: Bai, Guangji, et al.
Pubblicazione: (2024)
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
di: Yu, Yijiong, et al.
Pubblicazione: (2025)
di: Yu, Yijiong, et al.
Pubblicazione: (2025)
Variance Control via Weight Rescaling in LLM Pre-training
di: Owen, Louis, et al.
Pubblicazione: (2025)
di: Owen, Louis, et al.
Pubblicazione: (2025)
Dutch Metaphor Extraction from Cancer Patients' Interviews and Forum Data using LLMs and Human in the Loop
di: Han, Lifeng, et al.
Pubblicazione: (2025)
di: Han, Lifeng, et al.
Pubblicazione: (2025)
Evaluating NL2SQL via SQL2NL
di: Safarzadeh, Mohammadtaher, et al.
Pubblicazione: (2025)
di: Safarzadeh, Mohammadtaher, et al.
Pubblicazione: (2025)
FD-NL2SQL: Feedback-Driven Clinical NL2SQL that Improves with Use
di: Chowdhury, Suparno Roy, et al.
Pubblicazione: (2026)
di: Chowdhury, Suparno Roy, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Fietje: An open, efficient LLM for Dutch
di: Vanroy, Bram
Pubblicazione: (2024) -
GEITje 7B Ultra: A Conversational Model for Dutch
di: Vanroy, Bram
Pubblicazione: (2024) -
Everyone deserves their voice to be heard: Analyzing Predictive Gender Bias in ASR Models Applied to Dutch Speech Data
di: Raes, Rik, et al.
Pubblicazione: (2024) -
MTEB-NL and E5-NL: Embedding Benchmark and Models for Dutch
di: Banar, Nikolay, et al.
Pubblicazione: (2025) -
Biomedical Entity Linking for Dutch: Fine-tuning a Self-alignment BERT Model on an Automatically Generated Wikipedia Corpus
di: Hartendorp, Fons, et al.
Pubblicazione: (2024)