Got Compute, but No Data: Lessons From Post-training a Finnish LLM
Fuente:
arXiv
Guardado en:
| Autores principales: | Zosa, Elaine, Komulainen, Ville, Pyysalo, Sampo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Poro 34B and the Blessing of Multilinguality
por: Luukkonen, Risto, et al.
Publicado: (2024)
por: Luukkonen, Risto, et al.
Publicado: (2024)
Pretraining Finnish ModernBERTs
por: Reunamo, Akseli, et al.
Publicado: (2025)
por: Reunamo, Akseli, et al.
Publicado: (2025)
A Survey of Large Language Models for European Languages
por: Ali, Wazir, et al.
Publicado: (2024)
por: Ali, Wazir, et al.
Publicado: (2024)
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
por: Myntti, Amanda, et al.
Publicado: (2025)
por: Myntti, Amanda, et al.
Publicado: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
por: Kytöniemi, Joona, et al.
Publicado: (2025)
por: Kytöniemi, Joona, et al.
Publicado: (2025)
Scaling Data-Constrained Language Models
por: Muennighoff, Niklas, et al.
Publicado: (2023)
por: Muennighoff, Niklas, et al.
Publicado: (2023)
SemEval-2024 Shared Task 6: SHROOM, a Shared-task on Hallucinations and Related Observable Overgeneration Mistakes
por: Mickus, Timothee, et al.
Publicado: (2024)
por: Mickus, Timothee, et al.
Publicado: (2024)
Combining Qualitative and Computational Approaches for Literary Analysis of Finnish Novels
por: Ohman, Emily, et al.
Publicado: (2024)
por: Ohman, Emily, et al.
Publicado: (2024)
On the Impact of Calibration Data in Post-training Quantization and Pruning
por: Williams, Miles, et al.
Publicado: (2023)
por: Williams, Miles, et al.
Publicado: (2023)
Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
por: Nezhurina, Marianna, et al.
Publicado: (2025)
por: Nezhurina, Marianna, et al.
Publicado: (2025)
Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment
por: Zhang, Yizhuo, et al.
Publicado: (2025)
por: Zhang, Yizhuo, et al.
Publicado: (2025)
Asymmetric Conflict and Synergy in Post-training for LLM-based Multilingual Machine Translation
por: Zheng, Tong, et al.
Publicado: (2025)
por: Zheng, Tong, et al.
Publicado: (2025)
Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
por: Jing, Yi, et al.
Publicado: (2026)
por: Jing, Yi, et al.
Publicado: (2026)
On Predicting the Post-training Potential of Pre-trained LLMs
por: Li, Xiaoyuan, et al.
Publicado: (2026)
por: Li, Xiaoyuan, et al.
Publicado: (2026)
A New Massive Multilingual Dataset for High-Performance Language Technologies
por: de Gibert, Ona, et al.
Publicado: (2024)
por: de Gibert, Ona, et al.
Publicado: (2024)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
por: Ding, Bowen, et al.
Publicado: (2025)
por: Ding, Bowen, et al.
Publicado: (2025)
AIR: Post-training Data Selection for Reasoning via Attention Head Influence
por: Liu, Jinrui, et al.
Publicado: (2025)
por: Liu, Jinrui, et al.
Publicado: (2025)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
por: Finlayson, Matthew, et al.
Publicado: (2025)
por: Finlayson, Matthew, et al.
Publicado: (2025)
On Data Synthesis and Post-training for Visual Abstract Reasoning
por: Zhu, Ke, et al.
Publicado: (2025)
por: Zhu, Ke, et al.
Publicado: (2025)
LLM-based Triplet Extraction from Financial Reports
por: Wesslund, Dante, et al.
Publicado: (2026)
por: Wesslund, Dante, et al.
Publicado: (2026)
LLMs Got Rhythm? Hybrid Phonological Filtering for Greek Poetry Rhyme Detection and Generation
por: Chatzikyriakidis, Stergios, et al.
Publicado: (2026)
por: Chatzikyriakidis, Stergios, et al.
Publicado: (2026)
An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
por: Burchell, Laurie, et al.
Publicado: (2025)
por: Burchell, Laurie, et al.
Publicado: (2025)
DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training
por: Wang, Zhenting, et al.
Publicado: (2025)
por: Wang, Zhenting, et al.
Publicado: (2025)
Lessons from the Trenches on Reproducible Evaluation of Language Models
por: Biderman, Stella, et al.
Publicado: (2024)
por: Biderman, Stella, et al.
Publicado: (2024)
Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Training
por: Lacombe, Valentin, et al.
Publicado: (2026)
por: Lacombe, Valentin, et al.
Publicado: (2026)
T-RAG: Lessons from the LLM Trenches
por: Fatehkia, Masoomali, et al.
Publicado: (2024)
por: Fatehkia, Masoomali, et al.
Publicado: (2024)
Progress or Regress? Self-Improvement Reversal in Post-training
por: Wu, Ting, et al.
Publicado: (2024)
por: Wu, Ting, et al.
Publicado: (2024)
LLMs' morphological analyses of complex FST-generated Finnish words
por: Moisio, Anssi, et al.
Publicado: (2024)
por: Moisio, Anssi, et al.
Publicado: (2024)
Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
por: Qi, Jianyu, et al.
Publicado: (2025)
por: Qi, Jianyu, et al.
Publicado: (2025)
Best Practices and Lessons Learned on Synthetic Data
por: Liu, Ruibo, et al.
Publicado: (2024)
por: Liu, Ruibo, et al.
Publicado: (2024)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
por: Xu, Shanshan, et al.
Publicado: (2025)
por: Xu, Shanshan, et al.
Publicado: (2025)
RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
por: Jiang, Haoxiang, et al.
Publicado: (2026)
por: Jiang, Haoxiang, et al.
Publicado: (2026)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
por: Djuhera, Aladin, et al.
Publicado: (2025)
por: Djuhera, Aladin, et al.
Publicado: (2025)
Harnessing the Power of Multiple Minds: Lessons Learned from LLM Routing
por: Srivatsa, KV Aditya, et al.
Publicado: (2024)
por: Srivatsa, KV Aditya, et al.
Publicado: (2024)
PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
por: Liu, Jinhan, et al.
Publicado: (2026)
por: Liu, Jinhan, et al.
Publicado: (2026)
Boundary Suppression Asymmetry in Post-trained Assistants: Over-expansion as a Controllability Cost
por: Han, Jiarui
Publicado: (2026)
por: Han, Jiarui
Publicado: (2026)
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
por: Naous, Tarek, et al.
Publicado: (2025)
por: Naous, Tarek, et al.
Publicado: (2025)
I've Got 99 Problems But FLOPS Ain't One
por: Gherghescu, Alexandru M., et al.
Publicado: (2024)
por: Gherghescu, Alexandru M., et al.
Publicado: (2024)
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
por: Meeus, Matthieu, et al.
Publicado: (2024)
por: Meeus, Matthieu, et al.
Publicado: (2024)
Finnish SQuAD: A Simple Approach to Machine Translation of Span Annotations
por: Nuutinen, Emil, et al.
Publicado: (2025)
por: Nuutinen, Emil, et al.
Publicado: (2025)
Ejemplares similares
-
Poro 34B and the Blessing of Multilinguality
por: Luukkonen, Risto, et al.
Publicado: (2024) -
Pretraining Finnish ModernBERTs
por: Reunamo, Akseli, et al.
Publicado: (2025) -
A Survey of Large Language Models for European Languages
por: Ali, Wazir, et al.
Publicado: (2024) -
Register Always Matters: Analysis of LLM Pretraining Data Through the Lens of Language Variation
por: Myntti, Amanda, et al.
Publicado: (2025) -
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
por: Kytöniemi, Joona, et al.
Publicado: (2025)