Open-sci-ref-0.01: open and reproducible reference baselines for language model and dataset comparison
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nezhurina, Marianna, Franke, Jörg, Nakamura, Taishi, Carstensen, Timur, Ajroldi, Niccolò, Komulainen, Ville, Salinas, David, Jitsev, Jenia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
von: Cipolina-Kun, Lucia, et al.
Veröffentlicht: (2025)
von: Cipolina-Kun, Lucia, et al.
Veröffentlicht: (2025)
Learning in Compact Spaces with Approximately Normalized Transformer
von: Franke, Jörg K. H., et al.
Veröffentlicht: (2025)
von: Franke, Jörg K. H., et al.
Veröffentlicht: (2025)
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2024)
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2024)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025)
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
von: Nguyen, Huu, et al.
Veröffentlicht: (2025)
von: Nguyen, Huu, et al.
Veröffentlicht: (2025)
When, Where and Why to Average Weights?
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
von: Ajroldi, Niccolò, et al.
Veröffentlicht: (2025)
AudioToolAgent: An Agentic Framework for Audio-Language Models
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2025)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2025)
Got Compute, but No Data: Lessons From Post-training a Finnish LLM
von: Zosa, Elaine, et al.
Veröffentlicht: (2025)
von: Zosa, Elaine, et al.
Veröffentlicht: (2025)
Training Dynamics Impact Post-Training Quantization Robustness
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
von: Catalan-Tatjer, Albert, et al.
Veröffentlicht: (2025)
Resolving Discrepancies in Compute-Optimal Scaling of Language Models
von: Porian, Tomer, et al.
Veröffentlicht: (2024)
von: Porian, Tomer, et al.
Veröffentlicht: (2024)
Can you Finetune your Binoculars? Embedding Text Watermarks into the Weights of Large Language Models
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
von: Elhassan, Fay, et al.
Veröffentlicht: (2025)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
von: Islamov, Rustem, et al.
Veröffentlicht: (2024)
Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-size
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
von: Islamov, Rustem, et al.
Veröffentlicht: (2025)
Reproducible scaling laws for contrastive language-image learning
von: Cherti, Mehdi, et al.
Veröffentlicht: (2022)
von: Cherti, Mehdi, et al.
Veröffentlicht: (2022)
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
von: Udandarao, Vishaal, et al.
Veröffentlicht: (2025)
Inverse Deep Learning Ray Tracing for Heliostat Surface Prediction
von: Lewen, Jan, et al.
Veröffentlicht: (2024)
von: Lewen, Jan, et al.
Veröffentlicht: (2024)
Scalable heliostat surface predictions from focal spots: Sim-to-Real transfer of inverse Deep Learning Raytracing
von: Lewen, Jan, et al.
Veröffentlicht: (2025)
von: Lewen, Jan, et al.
Veröffentlicht: (2025)
Quickly Tuning Foundation Models for Image Segmentation
von: Das, Breenda, et al.
Veröffentlicht: (2025)
von: Das, Breenda, et al.
Veröffentlicht: (2025)
Frozen Layers: Memory-efficient Many-fidelity Hyperparameter Optimization
von: Carstensen, Timur, et al.
Veröffentlicht: (2025)
von: Carstensen, Timur, et al.
Veröffentlicht: (2025)
Lab ref: a handbook of recipes, reagents, and other reference tools for use at the bench / editor, Jane Roskams, Linda Rodgers
Scattering amplitudes of stable curves
von: Tevelev, Jenia
Veröffentlicht: (2020)
von: Tevelev, Jenia
Veröffentlicht: (2020)
B-HAR: an open-source baseline framework for in depth study of human activity recognition datasets and workflows
von: Demrozi, Florenc, et al.
Veröffentlicht: (2021)
von: Demrozi, Florenc, et al.
Veröffentlicht: (2021)
Deriving Hyperparameter Scaling Laws via Modern Optimization Theory
von: Shulgin, Egor, et al.
Veröffentlicht: (2026)
von: Shulgin, Egor, et al.
Veröffentlicht: (2026)
SelfAge: Personalized Facial Age Transformation Using Self-reference Images
von: Ito, Taishi, et al.
Veröffentlicht: (2025)
von: Ito, Taishi, et al.
Veröffentlicht: (2025)
TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
von: Moroshan, Vladyslav, et al.
Veröffentlicht: (2025)
von: Moroshan, Vladyslav, et al.
Veröffentlicht: (2025)
NicoleAdams-sci/ANHU_geneflow: ANHU_geneflow_1.0.0
von: Nicole
Veröffentlicht: (2025)
von: Nicole
Veröffentlicht: (2025)
31P-MRS reproducibility - datasets and code.
von: Svensen, Magnus
Veröffentlicht: (2025)
von: Svensen, Magnus
Veröffentlicht: (2025)
An open framework for archival, reproducible, and transparent science
von: Dasgupta, Sabar, et al.
Veröffentlicht: (2025)
von: Dasgupta, Sabar, et al.
Veröffentlicht: (2025)
Survival and growth of local and transplanted blue mussels (Mytilus trossulus, Lamark). / Jenia F. Yanick
von: Yanick, Jenia F
Veröffentlicht: (2003)
von: Yanick, Jenia F
Veröffentlicht: (2003)
On the Optimal Reasoning Length for RL-Trained Language Models
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
von: Wu, Yusong, et al.
Veröffentlicht: (2022)
PlotPick: AI-powered batch extraction of numerical data from scientific figures
von: Carstensen, Tommy
Veröffentlicht: (2026)
von: Carstensen, Tommy
Veröffentlicht: (2026)
Social Media in der Arbeitswelt
von: Carstensen, Tanja
Veröffentlicht: (2018)
von: Carstensen, Tanja
Veröffentlicht: (2018)
The Readers' Advisory Guide to Teen Literature
von: Carstensen, Angela
Veröffentlicht: (2018)
von: Carstensen, Angela
Veröffentlicht: (2018)
La maquila clandestina: el trabajo a domicilio informal en la Industria Textil y del Vestido en Puebla, México
von: Lisa Carstensen
Veröffentlicht: (2012)
von: Lisa Carstensen
Veröffentlicht: (2012)
Deformations of Kalck--Karmazyn algebras via Mirror Symmetry
von: Lekili, Yanki, et al.
Veröffentlicht: (2024)
von: Lekili, Yanki, et al.
Veröffentlicht: (2024)
Noncommutative resolution of $SU_C(2)$
von: Sink, Elias, et al.
Veröffentlicht: (2024)
von: Sink, Elias, et al.
Veröffentlicht: (2024)
Concept-Aware Batch Sampling Improves Language-Image Pretraining
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2025)
von: Ghosh, Adhiraj, et al.
Veröffentlicht: (2025)
MEDS-Tab: Automated tabularization and baseline methods for MEDS datasets
von: Oufattole, Nassim, et al.
Veröffentlicht: (2024)
von: Oufattole, Nassim, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Game Reasoning Arena: A Framework and Benchmark for Assessing Reasoning Capabilities of Large Language Models via Game Play
von: Cipolina-Kun, Lucia, et al.
Veröffentlicht: (2025) -
Learning in Compact Spaces with Approximately Normalized Transformer
von: Franke, Jörg K. H., et al.
Veröffentlicht: (2025) -
Alice in Wonderland: Simple Tasks Showing Complete Reasoning Breakdown in State-Of-the-Art Large Language Models
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2024) -
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
von: Nezhurina, Marianna, et al.
Veröffentlicht: (2025) -
MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources
von: Nguyen, Huu, et al.
Veröffentlicht: (2025)