The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Gienapp, Lukas, Schröder, Christopher, Schweter, Stefan, Akiki, Christopher, Schlatt, Ferdinand, Zimmermann, Arden, Genêt, Phillipe, Potthast, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SindBERT, the Sailor: Charting the Seas of Turkish NLP
por: Schmitt, Raphael, et al.
Publicado: (2025)
por: Schmitt, Raphael, et al.
Publicado: (2025)
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
por: Hoffmann, Michael, et al.
Publicado: (2025)
por: Hoffmann, Michael, et al.
Publicado: (2025)
German Text Embedding Clustering Benchmark
por: Wehrli, Silvan, et al.
Publicado: (2024)
por: Wehrli, Silvan, et al.
Publicado: (2024)
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
por: Akiki, Christopher, et al.
Publicado: (2023)
por: Akiki, Christopher, et al.
Publicado: (2023)
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
por: Gienapp, Lukas, et al.
Publicado: (2024)
por: Gienapp, Lukas, et al.
Publicado: (2024)
TITE Experiment Run Files
por: Schlatt, Ferdinand
Publicado: (2025)
por: Schlatt, Ferdinand
Publicado: (2025)
Historical German Text Normalization Using Type- and Token-Based Language Modeling
por: Ehrmanntraut, Anton
Publicado: (2024)
por: Ehrmanntraut, Anton
Publicado: (2024)
DETECT: Determining Ease and Textual Clarity of German Text Simplifications
por: Korobeynikova, Maria, et al.
Publicado: (2025)
por: Korobeynikova, Maria, et al.
Publicado: (2025)
Topic-Specific Classifiers are Better Relevance Judges than Prompted LLMs
por: Gienapp, Lukas, et al.
Publicado: (2025)
por: Gienapp, Lukas, et al.
Publicado: (2025)
The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
por: Kandpal, Nikhil, et al.
Publicado: (2025)
por: Kandpal, Nikhil, et al.
Publicado: (2025)
Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language Models
por: Schröder, Christopher, et al.
Publicado: (2024)
por: Schröder, Christopher, et al.
Publicado: (2024)
Ideology Prediction of German Political Texts
por: Schneider, Sinclair, et al.
Publicado: (2026)
por: Schneider, Sinclair, et al.
Publicado: (2026)
StylusAI: Stylistic Adaptation for Robust German Handwritten Text Generation
por: Riaz, Nauman, et al.
Publicado: (2024)
por: Riaz, Nauman, et al.
Publicado: (2024)
Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing
por: Darji, Harshil, et al.
Publicado: (2026)
por: Darji, Harshil, et al.
Publicado: (2026)
TL;DR Progress: Multi-faceted Literature Exploration in Text Summarization
por: Syed, Shahbaz, et al.
Publicado: (2024)
por: Syed, Shahbaz, et al.
Publicado: (2024)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
por: Waldis, Andreas, et al.
Publicado: (2024)
por: Waldis, Andreas, et al.
Publicado: (2024)
How Do Lexical Senses Correspond Between Spoken German and German Sign Language?
por: Çelikkol, Melis, et al.
Publicado: (2026)
por: Çelikkol, Melis, et al.
Publicado: (2026)
Evaluating Generative Ad Hoc Information Retrieval
por: Gienapp, Lukas, et al.
Publicado: (2023)
por: Gienapp, Lukas, et al.
Publicado: (2023)
Lightning IR: Straightforward Fine-tuning and Inference of Transformer-based Language Models for Information Retrieval
por: Schlatt, Ferdinand, et al.
Publicado: (2024)
por: Schlatt, Ferdinand, et al.
Publicado: (2024)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
por: Idrissi-Yaghir, Ahmad, et al.
Publicado: (2024)
por: Idrissi-Yaghir, Ahmad, et al.
Publicado: (2024)
CO-Fun: A German Dataset on Company Outsourcing in Fund Prospectuses for Named Entity Recognition and Relation Extraction
por: Foroutan, Neda, et al.
Publicado: (2024)
por: Foroutan, Neda, et al.
Publicado: (2024)
German Text Simplification: Finetuning Large Language Models with Semi-Synthetic Data
por: Klöser, Lars, et al.
Publicado: (2024)
por: Klöser, Lars, et al.
Publicado: (2024)
Segmentation and Processing of German Court Decisions from Open Legal Data
por: Darji, Harshil, et al.
Publicado: (2026)
por: Darji, Harshil, et al.
Publicado: (2026)
PolInterviews -- A Dataset of German Politician Public Broadcast Interviews
por: Birkenmaier, Lukas, et al.
Publicado: (2025)
por: Birkenmaier, Lukas, et al.
Publicado: (2025)
GottBERT: a pure German Language Model
por: Scheible, Raphael, et al.
Publicado: (2020)
por: Scheible, Raphael, et al.
Publicado: (2020)
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection
por: Herrlein, Janek, et al.
Publicado: (2024)
por: Herrlein, Janek, et al.
Publicado: (2024)
Investigating Counterclaims in Causality Extraction from Text
por: Hagen, Tim, et al.
Publicado: (2025)
por: Hagen, Tim, et al.
Publicado: (2025)
Adaptation and Evaluation of a German Sign Language Test
por: Haug, Tobias
Publicado: (2018)
por: Haug, Tobias
Publicado: (2018)
Large Language Model Data Generation for Enhanced Intent Recognition in German Speech
por: Rosin, Theresa Pekarek, et al.
Publicado: (2025)
por: Rosin, Theresa Pekarek, et al.
Publicado: (2025)
The clausal syntax of German Sign Language
por: Bross, Fabian
Publicado: (2020)
por: Bross, Fabian
Publicado: (2020)
German4All -- A Dataset and Model for Readability-Controlled Paraphrasing in German
por: Anschütz, Miriam, et al.
Publicado: (2025)
por: Anschütz, Miriam, et al.
Publicado: (2025)
Profiling German Text Simplification with Interpretable Model-Fingerprints
por: Klöser, Lars, et al.
Publicado: (2026)
por: Klöser, Lars, et al.
Publicado: (2026)
Ask, Retrieve, Summarize: A Modular Pipeline for Scientific Literature Summarization
por: Achkar, Pierre, et al.
Publicado: (2025)
por: Achkar, Pierre, et al.
Publicado: (2025)
Encyclopaedia of German diatheses
Publicado: (2023)
Publicado: (2023)
Classification of Human- and AI-Generated Texts for English, French, German, and Spanish
por: Schaaff, Kristina, et al.
Publicado: (2023)
por: Schaaff, Kristina, et al.
Publicado: (2023)
Large Language Models Discriminate Against Speakers of German Dialects
por: Bui, Minh Duc, et al.
Publicado: (2025)
por: Bui, Minh Duc, et al.
Publicado: (2025)
The Viability of Crowdsourcing for RAG Evaluation
por: Gienapp, Lukas, et al.
Publicado: (2025)
por: Gienapp, Lukas, et al.
Publicado: (2025)
A Teacher's Notebook: German.
Publicado: (1973)
Publicado: (1973)
Investigating the Effects of Sparse Attention on Cross-Encoders
por: Schlatt, Ferdinand, et al.
Publicado: (2023)
por: Schlatt, Ferdinand, et al.
Publicado: (2023)
Sentiment Analysis of German Sign Language Fairy Tales
por: Nunnari, Fabrizio, et al.
Publicado: (2026)
por: Nunnari, Fabrizio, et al.
Publicado: (2026)
Ejemplares similares
-
SindBERT, the Sailor: Charting the Seas of Turkish NLP
por: Schmitt, Raphael, et al.
Publicado: (2025) -
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
por: Hoffmann, Michael, et al.
Publicado: (2025) -
German Text Embedding Clustering Benchmark
por: Wehrli, Silvan, et al.
Publicado: (2024) -
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
por: Akiki, Christopher, et al.
Publicado: (2023) -
Learning Effective Representations for Retrieval Using Self-Distillation with Adaptive Relevance Margins
por: Gienapp, Lukas, et al.
Publicado: (2024)