Bundesrecht: An Open Library and Corpus for German Statutory Reference Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Darji, Harshil, Heckelmann, Martin, Kratsch, Christina, de Melo, Gerard |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Segmentation and Processing of German Court Decisions from Open Legal Data
by: Darji, Harshil, et al.
Published: (2026)
by: Darji, Harshil, et al.
Published: (2026)
Challenges and Considerations in Annotating Legal Data: A Comprehensive Overview
by: Darji, Harshil, et al.
Published: (2024)
by: Darji, Harshil, et al.
Published: (2024)
Investigating Sparsity in Recurrent Neural Networks
by: Darji, Harshil
Published: (2024)
by: Darji, Harshil
Published: (2024)
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
Statutory Construction and Interpretation for Artificial Intelligence
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
SliceMoE: Routing Embedding Slices Instead of Tokens for Fine-Grained and Balanced Transformer Scaling
by: Vejendla, Harshil
Published: (2025)
by: Vejendla, Harshil
Published: (2025)
Benchmarking Legal RAG: The Promise and Limits of AI Statutory Surveys
by: Afane, Mohamed, et al.
Published: (2026)
by: Afane, Mohamed, et al.
Published: (2026)
MDC-R: The Minecraft Dialogue Corpus with Reference
by: Madge, Chris, et al.
Published: (2025)
by: Madge, Chris, et al.
Published: (2025)
Bilingual BSARD: Extending Statutory Article Retrieval to Dutch
by: Lotfi, Ehsan, et al.
Published: (2024)
by: Lotfi, Ehsan, et al.
Published: (2024)
SwissGPC v1.0 -- The Swiss German Podcasts Corpus
by: Stucki, Samuel, et al.
Published: (2025)
by: Stucki, Samuel, et al.
Published: (2025)
A Cross-Lingual Statutory Article Retrieval Dataset for Taiwan Legal Studies
by: Wang, Yen-Hsiang, et al.
Published: (2024)
by: Wang, Yen-Hsiang, et al.
Published: (2024)
The German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
by: Gienapp, Lukas, et al.
Published: (2025)
by: Gienapp, Lukas, et al.
Published: (2025)
Transformer-Based Extraction of Statutory Definitions from the U.S. Code
by: Hosabettu, Arpana, et al.
Published: (2025)
by: Hosabettu, Arpana, et al.
Published: (2025)
QABISAR: Query-Article Bipartite Interactions for Statutory Article Retrieval
by: Santosh, T. Y. S. S., et al.
Published: (2024)
by: Santosh, T. Y. S. S., et al.
Published: (2024)
Connecting the Dots: What Graph-Based Text Representations Work Best for Text Classification Using Graph Neural Networks?
by: Bugueño, Margarita, et al.
Published: (2023)
by: Bugueño, Margarita, et al.
Published: (2023)
FOCUS: Effective Embedding Initialization for Monolingual Specialization of Multilingual Models
by: Dobler, Konstantin, et al.
Published: (2023)
by: Dobler, Konstantin, et al.
Published: (2023)
Rethinking Graph-Based Document Classification: Learning Data-Driven Structures Beyond Heuristic Approaches
by: Bugueño, Margarita, et al.
Published: (2025)
by: Bugueño, Margarita, et al.
Published: (2025)
PoTeC: A German Naturalistic Eye-tracking-while-reading Corpus
by: Jakobi, Deborah N., et al.
Published: (2024)
by: Jakobi, Deborah N., et al.
Published: (2024)
AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer Contracts
by: Braun, Daniel, et al.
Published: (2024)
by: Braun, Daniel, et al.
Published: (2024)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
by: Phatthiyaphaibun, Wannaphong, et al.
Published: (2025)
Developing an Open Conversational Speech Corpus for the Isan Language
by: Na-Thalang, Adisai, et al.
Published: (2025)
by: Na-Thalang, Adisai, et al.
Published: (2025)
DEPLAIN: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document Simplification
by: Stodden, Regina, et al.
Published: (2023)
by: Stodden, Regina, et al.
Published: (2023)
CEO: Corpus-based Open-Domain Event Ontology Induction
by: Xu, Nan, et al.
Published: (2023)
by: Xu, Nan, et al.
Published: (2023)
Language Adaptation on a Tight Academic Compute Budget: Tokenizer Swapping Works and Pure bfloat16 Is Enough
by: Dobler, Konstantin, et al.
Published: (2024)
by: Dobler, Konstantin, et al.
Published: (2024)
CuSINeS: Curriculum-driven Structure Induced Negative Sampling for Statutory Article Retrieval
by: Santosh, T. Y. S. S, et al.
Published: (2024)
by: Santosh, T. Y. S. S, et al.
Published: (2024)
Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
by: Soldaini, Luca, et al.
Published: (2024)
by: Soldaini, Luca, et al.
Published: (2024)
Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
by: Krsteski, Stefan, et al.
Published: (2025)
by: Krsteski, Stefan, et al.
Published: (2025)
ANHALTEN: Cross-Lingual Transfer for German Token-Level Reference-Free Hallucination Detection
by: Herrlein, Janek, et al.
Published: (2024)
by: Herrlein, Janek, et al.
Published: (2024)
Mitigate the Gap: Investigating Approaches for Improving Cross-Modal Alignment in CLIP
by: Eslami, Sedigheh, et al.
Published: (2024)
by: Eslami, Sedigheh, et al.
Published: (2024)
NextLevelBERT: Masked Language Modeling with Higher-Level Representations for Long Documents
by: Czinczoll, Tamara, et al.
Published: (2024)
by: Czinczoll, Tamara, et al.
Published: (2024)
InFact: Informativeness Alignment for Improved LLM Factuality
by: Cohen, Roi, et al.
Published: (2025)
by: Cohen, Roi, et al.
Published: (2025)
Pretrained LLMs Learn Multiple Types of Uncertainty
by: Cohen, Roi, et al.
Published: (2025)
by: Cohen, Roi, et al.
Published: (2025)
Token Distillation: Attention-aware Input Embeddings For New Tokens
by: Dobler, Konstantin, et al.
Published: (2025)
by: Dobler, Konstantin, et al.
Published: (2025)
CommitBench: A Benchmark for Commit Message Generation
by: Schall, Maximilian, et al.
Published: (2024)
by: Schall, Maximilian, et al.
Published: (2024)
The Open Proof Corpus: A Large-Scale Study of LLM-Generated Mathematical Proofs
by: Dekoninck, Jasper, et al.
Published: (2025)
by: Dekoninck, Jasper, et al.
Published: (2025)
OpenCSG Chinese Corpus: A Series of High-quality Chinese Datasets for LLM Training
by: Yu, Yijiong, et al.
Published: (2025)
by: Yu, Yijiong, et al.
Published: (2025)
Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Texts
by: Song, Seyoung, et al.
Published: (2025)
by: Song, Seyoung, et al.
Published: (2025)
MzansiText and MzansiLM: An Open Corpus and Decoder-Only Language Model for South African Languages
by: Lombard, Anri, et al.
Published: (2026)
by: Lombard, Anri, et al.
Published: (2026)
Query-Level Uncertainty in Large Language Models
by: Chen, Lihu, et al.
Published: (2025)
by: Chen, Lihu, et al.
Published: (2025)
Similar Items
-
Segmentation and Processing of German Court Decisions from Open Legal Data
by: Darji, Harshil, et al.
Published: (2026) -
Challenges and Considerations in Annotating Legal Data: A Comprehensive Overview
by: Darji, Harshil, et al.
Published: (2024) -
Investigating Sparsity in Recurrent Neural Networks
by: Darji, Harshil
Published: (2024) -
H1B-KV: Hybrid One-Bit Caches for Memory-Efficient Large Language Model Inference
by: Vejendla, Harshil
Published: (2025) -
Wave-PDE Nets: Trainable Wave-Equation Layers as an Alternative to Attention
by: Vejendla, Harshil
Published: (2025)