The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Gienapp, Lukas, Schröder, Christopher, Schweter, Stefan, Akiki, Christopher, Schlatt, Ferdinand, Zimmermann, Arden, Genêt, Phillipe, Potthast, Martin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SindBERT, the Sailor: Charting the Seas of Turkish NLP
di: Schmitt, Raphael, et al.
Pubblicazione: (2025)
di: Schmitt, Raphael, et al.
Pubblicazione: (2025)
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
German Text Embedding Clustering Benchmark
di: Wehrli, Silvan, et al.
Pubblicazione: (2024)
di: Wehrli, Silvan, et al.
Pubblicazione: (2024)
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
di: Akiki, Christopher, et al.
Pubblicazione: (2023)
di: Akiki, Christopher, et al.
Pubblicazione: (2023)
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
di: Gienapp, Lukas, et al.
Pubblicazione: (2024)
di: Gienapp, Lukas, et al.
Pubblicazione: (2024)
TITE Experiment Run Files
di: Schlatt, Ferdinand
Pubblicazione: (2025)
di: Schlatt, Ferdinand
Pubblicazione: (2025)
Historical German Text Normalization Using Type- and Token-Based Language Modeling
di: Ehrmanntraut, Anton
Pubblicazione: (2024)
di: Ehrmanntraut, Anton
Pubblicazione: (2024)
DETECT: Determining Ease and Textual Clarity of German Text Simplifications
di: Korobeynikova, Maria, et al.
Pubblicazione: (2025)
di: Korobeynikova, Maria, et al.
Pubblicazione: (2025)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
di: Gienapp, Lukas, et al.
Pubblicazione: (2025)
di: Gienapp, Lukas, et al.
Pubblicazione: (2025)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
di: Schröder, Christopher, et al.
Pubblicazione: (2024)
Ideology Prediction of German Political Texts
di: Schneider, Sinclair, et al.
Pubblicazione: (2026)
di: Schneider, Sinclair, et al.
Pubblicazione: (2026)
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
di: Riaz, Nauman, et al.
Pubblicazione: (2024)
di: Riaz, Nauman, et al.
Pubblicazione: (2024)
Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing
di: Darji, Harshil, et al.
Pubblicazione: (2026)
di: Darji, Harshil, et al.
Pubblicazione: (2026)
TL;DR Progress: Multi-faceted Literature Exploration in Text Summarization
di: Syed, Shahbaz, et al.
Pubblicazione: (2024)
di: Syed, Shahbaz, et al.
Pubblicazione: (2024)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
How Do Lexical Senses Correspond Between Spoken German and German Sign Language?
di: Çelikkol, Melis, et al.
Pubblicazione: (2026)
di: Çelikkol, Melis, et al.
Pubblicazione: (2026)
Evaluating Generative Ad Hoc Information Retrieval
di: Gienapp, Lukas, et al.
Pubblicazione: (2023)
di: Gienapp, Lukas, et al.
Pubblicazione: (2023)
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
di: Schlatt, Ferdinand, et al.
Pubblicazione: (2024)
di: Schlatt, Ferdinand, et al.
Pubblicazione: (2024)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction
di: Foroutan, Neda, et al.
Pubblicazione: (2024)
di: Foroutan, Neda, et al.
Pubblicazione: (2024)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
di: Klöser, Lars, et al.
Pubblicazione: (2024)
di: Klöser, Lars, et al.
Pubblicazione: (2024)
Segmentation and Processing of German Court Decisions from Open Legal Data
di: Darji, Harshil, et al.
Pubblicazione: (2026)
di: Darji, Harshil, et al.
Pubblicazione: (2026)
PolInterviews -- A Dataset of German Politician Public Broadcast Interviews
di: Birkenmaier, Lukas, et al.
Pubblicazione: (2025)
di: Birkenmaier, Lukas, et al.
Pubblicazione: (2025)
GottBERT: a pure German Language Model
di: Scheible, Raphael, et al.
Pubblicazione: (2020)
di: Scheible, Raphael, et al.
Pubblicazione: (2020)
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection
di: Herrlein, Janek, et al.
Pubblicazione: (2024)
di: Herrlein, Janek, et al.
Pubblicazione: (2024)
Investigating Counterclaims in Causality Extraction from Text
di: Hagen, Tim, et al.
Pubblicazione: (2025)
di: Hagen, Tim, et al.
Pubblicazione: (2025)
Adaptation and Evaluation of a German Sign Language Test
di: Haug, Tobias
Pubblicazione: (2018)
di: Haug, Tobias
Pubblicazione: (2018)
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
di: Rosin, Theresa Pekarek, et al.
Pubblicazione: (2025)
The clausal syntax of German Sign Language
di: Bross, Fabian
Pubblicazione: (2020)
di: Bross, Fabian
Pubblicazione: (2020)
German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
di: Anschütz, Miriam, et al.
Pubblicazione: (2025)
di: Anschütz, Miriam, et al.
Pubblicazione: (2025)
Profiling German Text Simplification with Interpretable Model-Fingerprints
di: Klöser, Lars, et al.
Pubblicazione: (2026)
di: Klöser, Lars, et al.
Pubblicazione: (2026)
Ask, Retrieve, Summarize: A Modular Pipeline for Scientific Literature Summarization
di: Achkar, Pierre, et al.
Pubblicazione: (2025)
di: Achkar, Pierre, et al.
Pubblicazione: (2025)
Encyclopaedia of German diatheses
Pubblicazione: (2023)
Pubblicazione: (2023)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
di: Schaaff, Kristina, et al.
Pubblicazione: (2023)
di: Schaaff, Kristina, et al.
Pubblicazione: (2023)
Large Language Models Discriminate Against Speakers of German Dialects
di: Bui, Minh Duc, et al.
Pubblicazione: (2025)
di: Bui, Minh Duc, et al.
Pubblicazione: (2025)
The Viability of Crowdsourcing for RAG Evaluation
di: Gienapp, Lukas, et al.
Pubblicazione: (2025)
di: Gienapp, Lukas, et al.
Pubblicazione: (2025)
A Teacher's Notebook: German.
Pubblicazione: (1973)
Pubblicazione: (1973)
Investigating the Effects of Sparse Attention on Cross-Encoders
di: Schlatt, Ferdinand, et al.
Pubblicazione: (2023)
di: Schlatt, Ferdinand, et al.
Pubblicazione: (2023)
Sentiment Analysis of German Sign Language Fairy Tales
di: Nunnari, Fabrizio, et al.
Pubblicazione: (2026)
di: Nunnari, Fabrizio, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SindBERT, the Sailor: Charting the Seas of Turkish NLP
di: Schmitt, Raphael, et al.
Pubblicazione: (2025) -
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
di: Hoffmann, Michael, et al.
Pubblicazione: (2025) -
German Text Embedding Clustering Benchmark
di: Wehrli, Silvan, et al.
Pubblicazione: (2024) -
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
di: Akiki, Christopher, et al.
Pubblicazione: (2023) -
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
di: Gienapp, Lukas, et al.
Pubblicazione: (2024)