The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Fuente:
arXiv
Salvato in:
| Autori principali: | Kandpal, Nikhil, Lester, Brian, Raffel, Colin, Majstorovic, Sebastian, Biderman, Stella, Abbasi, Baber, Soldaini, Luca, Shippole, Enrico, Cooper, A. Feder, Skowron, Aviya, Kirchenbauer, John, Longpre, Shayne, Sutawika, Lintang, Albalak, Alon, Xu, Zhenlin, Penedo, Guilherme, Allal, Loubna Ben, Bakouch, Elie, Pressman, John David, Fan, Honglu, Stander, Dashiell, Song, Guangyu, Gokaslan, Aaron, Goldstein, Tom, Bartoldson, Brian R., Kailkhura, Bhavya, Murray, Tyler |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Grokking Group Multiplication with Cosets
di: Stander, Dashiell, et al.
Pubblicazione: (2023)
di: Stander, Dashiell, et al.
Pubblicazione: (2023)
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
di: Wang, Zeyu, et al.
Pubblicazione: (2025)
di: Wang, Zeyu, et al.
Pubblicazione: (2025)
Position: The Most Expensive Part of an LLM should be its Training Data
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025)
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
di: Bartoldson, Brian R., et al.
Pubblicazione: (2024)
di: Bartoldson, Brian R., et al.
Pubblicazione: (2024)
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)
Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
di: Hägele, Alexander, et al.
Pubblicazione: (2024)
Multi-Token Prediction via Self-Distillation
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
di: Kirchenbauer, John, et al.
Pubblicazione: (2026)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
di: Geiping, Jonas, et al.
Pubblicazione: (2025)
di: Geiping, Jonas, et al.
Pubblicazione: (2025)
Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
di: McDonald, Tavish, et al.
Pubblicazione: (2025)
di: McDonald, Tavish, et al.
Pubblicazione: (2025)
YaRN: Efficient Context Window Extension of Large Language Models
di: Peng, Bowen, et al.
Pubblicazione: (2023)
di: Peng, Bowen, et al.
Pubblicazione: (2023)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution
di: Liu, Fengyuan, et al.
Pubblicazione: (2024)
di: Liu, Fengyuan, et al.
Pubblicazione: (2024)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
di: McLeish, Sean, et al.
Pubblicazione: (2025)
di: McLeish, Sean, et al.
Pubblicazione: (2025)
Transformers Can Do Arithmetic with the Right Embeddings
di: McLeish, Sean, et al.
Pubblicazione: (2024)
di: McLeish, Sean, et al.
Pubblicazione: (2024)
Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
di: Zheng, Haizhong, et al.
Pubblicazione: (2025)
di: Zheng, Haizhong, et al.
Pubblicazione: (2025)
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion
di: Christopher, Jacob K, et al.
Pubblicazione: (2024)
di: Christopher, Jacob K, et al.
Pubblicazione: (2024)
ELFS: Label-Free Coreset Selection with Proxy Training Dynamics
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
di: Zheng, Haizhong, et al.
Pubblicazione: (2024)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
di: Cai, Zikui, et al.
Pubblicazione: (2025)
di: Cai, Zikui, et al.
Pubblicazione: (2025)
Enhancing Training Data Attribution with Representational Optimization
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
di: Sun, Weiwei, et al.
Pubblicazione: (2025)
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
di: Penedo, Guilherme, et al.
Pubblicazione: (2024)
di: Penedo, Guilherme, et al.
Pubblicazione: (2024)
The Responsible Foundation Model Development Cheatsheet: A Review of Tools & Resources
di: Longpre, Shayne, et al.
Pubblicazione: (2024)
di: Longpre, Shayne, et al.
Pubblicazione: (2024)
Self-Directed Synthetic Dialogues and Revisions Technical Report
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
di: Lambert, Nathan, et al.
Pubblicazione: (2024)
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
di: Bradshaw, Louis, et al.
Pubblicazione: (2025)
di: Bradshaw, Louis, et al.
Pubblicazione: (2025)
STAR-1: Safer Alignment of Reasoning LLMs with 1K Data
di: Wang, Zijun, et al.
Pubblicazione: (2025)
di: Wang, Zijun, et al.
Pubblicazione: (2025)
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
di: Niklaus, Joel, et al.
Pubblicazione: (2026)
di: Niklaus, Joel, et al.
Pubblicazione: (2026)
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
di: Allal, Loubna Ben, et al.
Pubblicazione: (2025)
di: Allal, Loubna Ben, et al.
Pubblicazione: (2025)
The German‐Soviet War: Combat, Occupation, and Legacies by JeffRutherford and RobertvonMaier, eds. Ithaca: Cornell University Press, 2025. 600 pp. $59.95. ISBN 978‐1‐5017‐8108‐7
di: Vojin Majstorovic
Pubblicazione: (2026)
di: Vojin Majstorovic
Pubblicazione: (2026)
Literature of the americas in the making: U.S. writers and translation in sur, 1931-1944
di: Gorica Majstorovic
Pubblicazione: (2013)
di: Gorica Majstorovic
Pubblicazione: (2013)
EL YESO DE ZANCOLLI EN EL TRATAMIENTO DE LAS LESIONES TRAUMATICAS DE LA MANO EN NIÑOS
di: Dashiell Cañizares Betancourt
Pubblicazione: (2006)
di: Dashiell Cañizares Betancourt
Pubblicazione: (2006)
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training
di: Bartoldson, Brian, et al.
Pubblicazione: (2025)
di: Bartoldson, Brian, et al.
Pubblicazione: (2025)
Ingeniería del software : un enfoque pr ctico / Roger S. Pressman ; traductor Jesús Elmer Murrieta Murrieta, Eloy Pineda Rojas, Víctor Campos Olguín
di: Pressman, Roger S
Pubblicazione: (2005)
di: Pressman, Roger S
Pubblicazione: (2005)
Ingeniería del software : un enfoque pr ctico / Roger S. Pressman ; traductor, Víctor Campos Olguín, Javíer Enríquez Brito
di: Pressman, Roger S
Pubblicazione: (2010)
di: Pressman, Roger S
Pubblicazione: (2010)
Israeli Unilateralism and Israeli—Palestinian Relations, 2001-2006 / Jeremy Pressman
di: Pressman, Jeremy
Pubblicazione: (2001)
di: Pressman, Jeremy
Pubblicazione: (2001)
Ingeniería del software : un enfoque pr ctico / Roger S. Pressman ; traductor José María Troya, Luis Hern ndez Yañez
di: Pressman, Roger S
Pubblicazione: (1988)
di: Pressman, Roger S
Pubblicazione: (1988)
Software engineering; a practitioner's approach / Roger S. Pressman
di: Pressman, Roger S
Pubblicazione: (1982)
di: Pressman, Roger S
Pubblicazione: (1982)
Israeli Unilateralism and Israeli-Palestina Relations, 2001-2006
di: Pressman, Jeremy
Pubblicazione: (2001)
di: Pressman, Jeremy
Pubblicazione: (2001)
La clase media en países latinoamericanos
di: Steven Pressman
Pubblicazione: (2011)
di: Steven Pressman
Pubblicazione: (2011)
ENTERPRISE ARCHITECTURE AS AN APPROACH TO THE DEVELOPMENT OF INFORMATION SYSTEMS
di: Milosav N. Majstorović
Pubblicazione: (2018)
di: Milosav N. Majstorović
Pubblicazione: (2018)
BUSINESS AND IT ALIGNMENT
di: Milosav N. Majstorović
Pubblicazione: (2016)
di: Milosav N. Majstorović
Pubblicazione: (2016)
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
di: Yue, Xiang, et al.
Pubblicazione: (2024)
di: Yue, Xiang, et al.
Pubblicazione: (2024)
Driving AR Experience to Purchase Intention: Examining the Role of Users' Immersiveness, Perceived Innovativeness and Engagement in Virtual Try‐On
di: Ruturaj Baber, et al.
Pubblicazione: (2026)
di: Ruturaj Baber, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Grokking Group Multiplication with Cosets
di: Stander, Dashiell, et al.
Pubblicazione: (2023) -
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness
di: Wang, Zeyu, et al.
Pubblicazione: (2025) -
Position: The Most Expensive Part of an LLM should be its Training Data
di: Kandpal, Nikhil, et al.
Pubblicazione: (2025) -
Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies
di: Bartoldson, Brian R., et al.
Pubblicazione: (2024) -
Gained in Translation: Privileged Pairwise Judges Enhance Multilingual Reasoning
di: Sutawika, Lintang, et al.
Pubblicazione: (2026)