Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Ali, Mehdi, Fromm, Michael, Thellmann, Klaudia, Ebert, Jan, Weber, Alexander Arno, Rutmann, Richard, Jain, Charvi, Lübbering, Max, Steinigen, Daniel, Leveling, Johannes, Klug, Katrin, Buschhoff, Jasper Schulze, Jurkschat, Lena, Abdelwahab, Hammam, Stein, Benny Jörg, Sylla, Karl-Heinz, Denisov, Pavel, Brandizzi, Nicolo', Saleem, Qasid, Bhowmick, Anirban, Helmer, Lennard, John, Chelsea, Suarez, Pedro Ortiz, Ostendorff, Malte, Jude, Alex, Manjunath, Lalith, Weinbach, Samuel, Penke, Carolin, Filatov, Oleg, Barth, Fabio, Mirza, Paramita, Weber, Lucas, Wendler, Ines, Sifa, Rafet, Küch, Fabian, Herten, Andreas, Jäkel, René, Rehm, Georg, Kesselheim, Stefan, Köhler, Joachim, Flores-Herr, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Tokenizer Choice For LLM Training: Negligible or Crucial?
di: Ali, Mehdi, et al.
Pubblicazione: (2023)
di: Ali, Mehdi, et al.
Pubblicazione: (2023)
Teuken Bidikay, Revista Latinoamericana de Investigación en Organizaciones Ambiente y Sociedad
Pubblicazione: (2019)
Pubblicazione: (2019)
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
di: Lübbering, Max, et al.
Pubblicazione: (2026)
di: Lübbering, Max, et al.
Pubblicazione: (2026)
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
di: Penke, Carolin, et al.
Pubblicazione: (2025)
di: Penke, Carolin, et al.
Pubblicazione: (2025)
Data Processing for the OpenGPT-X Model Family
di: Brandizzi, Nicolo', et al.
Pubblicazione: (2024)
di: Brandizzi, Nicolo', et al.
Pubblicazione: (2024)
Quantum Computing from Hopfield Nets
di: Bauckhage, Christian, et al.
Pubblicazione: (2025)
di: Bauckhage, Christian, et al.
Pubblicazione: (2025)
Performance and Power: Systematic Evaluation of AI Workloads on Accelerators with CARAML
di: John, Chelsea Maria, et al.
Pubblicazione: (2024)
di: John, Chelsea Maria, et al.
Pubblicazione: (2024)
Towards Multilingual LLM Evaluation for European Languages
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2024)
SynCED-EnDe 2025: A Synthetic and Curated English - German Dataset for Critical Error Detection in Machine Translation
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
Towards Reliable Machine Translation: Scaling LLMs for Critical Error Detection and Safety
di: Chopra, Muskaan, et al.
Pubblicazione: (2026)
di: Chopra, Muskaan, et al.
Pubblicazione: (2026)
LASER: Stratified Selective Sampling for Instruction Tuning with Dedicated Scoring Strategy
di: Mirza, Paramita, et al.
Pubblicazione: (2025)
di: Mirza, Paramita, et al.
Pubblicazione: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
Evaluation von Verwaltungsmodernisierung
di: Buschhoff, Christian
Pubblicazione: (2020)
di: Buschhoff, Christian
Pubblicazione: (2020)
Time Transfer: On Optimal Learning Rate and Batch Size In The Infinite Data Limit
di: Filatov, Oleg, et al.
Pubblicazione: (2024)
di: Filatov, Oleg, et al.
Pubblicazione: (2024)
Memory and Bandwidth are All You Need for Fully Sharded Data Parallel
di: Wang, Jiangtao, et al.
Pubblicazione: (2025)
di: Wang, Jiangtao, et al.
Pubblicazione: (2025)
Optimal Scaling Needs Optimal Norm
di: Filatov, Oleg, et al.
Pubblicazione: (2025)
di: Filatov, Oleg, et al.
Pubblicazione: (2025)
Model-agnostic Body Part Relevance Assessment for Pedestrian Detection
di: Günder, Maurice, et al.
Pubblicazione: (2023)
di: Günder, Maurice, et al.
Pubblicazione: (2023)
[Vision Paper] PRObot: Enhancing Patient-Reported Outcome Measures for Diabetic Retinopathy using Chatbots and Generative AI
di: Pielka, Maren, et al.
Pubblicazione: (2024)
di: Pielka, Maren, et al.
Pubblicazione: (2024)
Pointer-Guided Pre-Training: Infusing Large Language Models with Paragraph-Level Contextual Awareness
di: Hillebrand, Lars, et al.
Pubblicazione: (2024)
di: Hillebrand, Lars, et al.
Pubblicazione: (2024)
Reasoning LLMs in the Medical Domain: A Literature Survey
di: Berger, Armin, et al.
Pubblicazione: (2025)
di: Berger, Armin, et al.
Pubblicazione: (2025)
Interpretable Topic Extraction and Word Embedding Learning using row-stochastic DEDICOM
di: Hillebrand, Lars, et al.
Pubblicazione: (2025)
di: Hillebrand, Lars, et al.
Pubblicazione: (2025)
How Small Can You Go? Compact Language Models for On-Device Critical Error Detection in Machine Translation
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening
di: Chopra, Muskaan, et al.
Pubblicazione: (2026)
di: Chopra, Muskaan, et al.
Pubblicazione: (2026)
Towards More Human-like AI Communication: A Review of Emergent Communication Research
di: Brandizzi, Nicolo'
Pubblicazione: (2023)
di: Brandizzi, Nicolo'
Pubblicazione: (2023)
Hacia donde vamos en la educación de salud bucal en Latinoamérica
di: Daniel Brandizzi
Pubblicazione: (2023)
di: Daniel Brandizzi
Pubblicazione: (2023)
History Rhymes: Macro-Contextual Retrieval for Robust Financial Forecasting
di: Khanna, Sarthak, et al.
Pubblicazione: (2025)
di: Khanna, Sarthak, et al.
Pubblicazione: (2025)
Generalizing Abstention for Noise-Robust Learning in Medical Image Segmentation
di: Moustafa, Wesam, et al.
Pubblicazione: (2026)
di: Moustafa, Wesam, et al.
Pubblicazione: (2026)
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
di: Berghaus, David, et al.
Pubblicazione: (2025)
di: Berghaus, David, et al.
Pubblicazione: (2025)
Can Continual Pre-training Bridge the Performance Gap between General-purpose and Specialized Language Models in the Medical Domain?
di: Doll, Niclas, et al.
Pubblicazione: (2026)
di: Doll, Niclas, et al.
Pubblicazione: (2026)
Liberdade social e socialização do mercado
di: Hannes Kuch
Pubblicazione: (2018)
di: Hannes Kuch
Pubblicazione: (2018)
¿Quiénes se benefician del turismo en Cayos Cochinos, Honduras?
di: Sebastian Kuch
Pubblicazione: (2015)
di: Sebastian Kuch
Pubblicazione: (2015)
A Survey on Current Trends and Recent Advances in Text Anonymization
di: Deußer, Tobias, et al.
Pubblicazione: (2025)
di: Deußer, Tobias, et al.
Pubblicazione: (2025)
Towards Unified Multimodal Financial Forecasting: Integrating Sentiment Embeddings and Market Indicators via Cross-Modal Attention
di: Khanna, Sarthak, et al.
Pubblicazione: (2025)
di: Khanna, Sarthak, et al.
Pubblicazione: (2025)
When the Whole Is Less Than the Sum of Its Parts: Structural Coupling in Education
di: Raf Vanderstraeten, et al.
Pubblicazione: (2026)
di: Raf Vanderstraeten, et al.
Pubblicazione: (2026)
Stärkung des Selbstbildes durch ein Unterrichtsprojekt rund um ein Bilderbuch Konzeptentwicklung zu "Darius Farbtupf" für 7- bis 9-jährige Unterstufenkinder
di: Weber, Barbara
Pubblicazione: (2014)
di: Weber, Barbara
Pubblicazione: (2014)
Is continuous CoT better suited for multi-lingual reasoning?
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
di: Bashir, Ali Hamza, et al.
Pubblicazione: (2026)
From Retinal Pixels to Patients: Evolution of Deep Learning Research in Diabetic Retinopathy Screening
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
di: Chopra, Muskaan, et al.
Pubblicazione: (2025)
Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite
di: Thellmann, Klaudia, et al.
Pubblicazione: (2026)
di: Thellmann, Klaudia, et al.
Pubblicazione: (2026)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
di: Weber, Alexander Arno, et al.
Pubblicazione: (2024)
di: Weber, Alexander Arno, et al.
Pubblicazione: (2024)
Unreachable, Inescapable: Sustainable Development as Normative Camouflage in EU–MERCOSUR Trade
di: Asha Herten‐Crabb
Pubblicazione: (2026)
di: Asha Herten‐Crabb
Pubblicazione: (2026)
Documenti analoghi
-
Tokenizer Choice For LLM Training: Negligible or Crucial?
di: Ali, Mehdi, et al.
Pubblicazione: (2023) -
Teuken Bidikay, Revista Latinoamericana de Investigación en Organizaciones Ambiente y Sociedad
Pubblicazione: (2019) -
Modalities, a PyTorch-native Framework For Large-scale LLM Training and Research
di: Lübbering, Max, et al.
Pubblicazione: (2026) -
Training LLMs on HPC Systems: Best Practices from the OpenGPT-X Project
di: Penke, Carolin, et al.
Pubblicazione: (2025) -
Data Processing for the OpenGPT-X Model Family
di: Brandizzi, Nicolo', et al.
Pubblicazione: (2024)