Beyond Public Access in LLM Pre-Training Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Rosenblat, Sruly, O'Reilly, Tim, Strauss, Ilan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Attribution Crisis in LLM Search Results
por: Strauss, Ilan, et al.
Publicado: (2025)
por: Strauss, Ilan, et al.
Publicado: (2025)
Real-World Gaps in AI Governance Research
por: Strauss, Ilan, et al.
Publicado: (2025)
por: Strauss, Ilan, et al.
Publicado: (2025)
LLM-Supported Natural Language to Bash Translation
por: Westenfelder, Finnian, et al.
Publicado: (2025)
por: Westenfelder, Finnian, et al.
Publicado: (2025)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
por: O'Reilly, Cliff, et al.
Publicado: (2025)
por: O'Reilly, Cliff, et al.
Publicado: (2025)
Experiments in News Bias Detection with Pre-Trained Neural Transformers
por: Menzner, Tim, et al.
Publicado: (2024)
por: Menzner, Tim, et al.
Publicado: (2024)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
por: Li, Shujie, et al.
Publicado: (2024)
por: Li, Shujie, et al.
Publicado: (2024)
Rethinking Reflection in Pre-Training
por: AI, Essential, et al.
Publicado: (2025)
por: AI, Essential, et al.
Publicado: (2025)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
por: Baroian, Andrei, et al.
Publicado: (2025)
por: Baroian, Andrei, et al.
Publicado: (2025)
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
por: Flor, Michael, et al.
Publicado: (2025)
por: Flor, Michael, et al.
Publicado: (2025)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
por: Jiang, Changhao, et al.
Publicado: (2025)
por: Jiang, Changhao, et al.
Publicado: (2025)
Reinforcement Learning on Pre-Training Data
por: Li, Siheng, et al.
Publicado: (2025)
por: Li, Siheng, et al.
Publicado: (2025)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
por: Fujii, Kazuki, et al.
Publicado: (2024)
por: Fujii, Kazuki, et al.
Publicado: (2024)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
por: Du, Wenyu, et al.
Publicado: (2024)
por: Du, Wenyu, et al.
Publicado: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
por: Weller, Orion, et al.
Publicado: (2023)
por: Weller, Orion, et al.
Publicado: (2023)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
por: Ceron, Tanise, et al.
Publicado: (2025)
por: Ceron, Tanise, et al.
Publicado: (2025)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
por: Masoudian, Shahed, et al.
Publicado: (2025)
por: Masoudian, Shahed, et al.
Publicado: (2025)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
por: Zheng, Weihua, et al.
Publicado: (2026)
por: Zheng, Weihua, et al.
Publicado: (2026)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
por: Shivagunde, Namrata, et al.
Publicado: (2026)
por: Shivagunde, Namrata, et al.
Publicado: (2026)
Unraveling Emotions with Pre-Trained Models
por: Pajón-Sanmartín, Alejandro, et al.
Publicado: (2025)
por: Pajón-Sanmartín, Alejandro, et al.
Publicado: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
por: Shen, Zhiqiang, et al.
Publicado: (2023)
por: Shen, Zhiqiang, et al.
Publicado: (2023)
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
por: Chung, Philip, et al.
Publicado: (2024)
por: Chung, Philip, et al.
Publicado: (2024)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
por: Tian, Junfeng, et al.
Publicado: (2024)
por: Tian, Junfeng, et al.
Publicado: (2024)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
por: Zhuo, Le, et al.
Publicado: (2024)
por: Zhuo, Le, et al.
Publicado: (2024)
Evolution of Concepts in Language Model Pre-Training
por: Ge, Xuyang, et al.
Publicado: (2025)
por: Ge, Xuyang, et al.
Publicado: (2025)
Beyond Memorization: The Challenge of Random Memory Access in Language Models
por: Zhu, Tongyao, et al.
Publicado: (2024)
por: Zhu, Tongyao, et al.
Publicado: (2024)
Scaling LLM Pre-training with Vocabulary Curriculum
por: Yu, Fangyuan
Publicado: (2025)
por: Yu, Fangyuan
Publicado: (2025)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
por: Xu, Yuanjian, et al.
Publicado: (2026)
por: Xu, Yuanjian, et al.
Publicado: (2026)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
por: Chung, Yi-Ling, et al.
Publicado: (2025)
por: Chung, Yi-Ling, et al.
Publicado: (2025)
DEFT: Data Efficient Fine-Tuning for Pre-Trained Language Models via Unsupervised Core-Set Selection
por: Das, Devleena, et al.
Publicado: (2023)
por: Das, Devleena, et al.
Publicado: (2023)
Revealing the Inherent Instructability of Pre-Trained Language Models
por: An, Seokhyun, et al.
Publicado: (2024)
por: An, Seokhyun, et al.
Publicado: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
por: Kang, Feiyang, et al.
Publicado: (2024)
por: Kang, Feiyang, et al.
Publicado: (2024)
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
por: Matta, Shiho, et al.
Publicado: (2024)
por: Matta, Shiho, et al.
Publicado: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
Pre-Trained Language Models for Keyphrase Prediction: A Review
por: Umair, Muhammad, et al.
Publicado: (2024)
por: Umair, Muhammad, et al.
Publicado: (2024)
Characterizing Stereotypical Bias from Privacy-preserving Pre-Training
por: Arnold, Stefan, et al.
Publicado: (2024)
por: Arnold, Stefan, et al.
Publicado: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
por: Qiu, Zeju, et al.
Publicado: (2025)
por: Qiu, Zeju, et al.
Publicado: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
por: Kim, Eunsu, et al.
Publicado: (2024)
por: Kim, Eunsu, et al.
Publicado: (2024)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
por: Rosenbaum, Andy, et al.
Publicado: (2026)
por: Rosenbaum, Andy, et al.
Publicado: (2026)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
por: Djuhera, Aladin, et al.
Publicado: (2025)
por: Djuhera, Aladin, et al.
Publicado: (2025)
Ejemplares similares
-
The Attribution Crisis in LLM Search Results
por: Strauss, Ilan, et al.
Publicado: (2025) -
Real-World Gaps in AI Governance Research
por: Strauss, Ilan, et al.
Publicado: (2025) -
LLM-Supported Natural Language to Bash Translation
por: Westenfelder, Finnian, et al.
Publicado: (2025) -
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
por: O'Reilly, Cliff, et al.
Publicado: (2025) -
Experiments in News Bias Detection with Pre-Trained Neural Transformers
por: Menzner, Tim, et al.
Publicado: (2024)