Beyond Public Access in LLM Pre-Training Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Rosenblat, Sruly, O'Reilly, Tim, Strauss, Ilan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Attribution Crisis in LLM Search Results
di: Strauss, Ilan, et al.
Pubblicazione: (2025)
di: Strauss, Ilan, et al.
Pubblicazione: (2025)
Real-World Gaps in AI Governance Research
di: Strauss, Ilan, et al.
Pubblicazione: (2025)
di: Strauss, Ilan, et al.
Pubblicazione: (2025)
LLM-Supported Natural Language to Bash Translation
di: Westenfelder, Finnian, et al.
Pubblicazione: (2025)
di: Westenfelder, Finnian, et al.
Pubblicazione: (2025)
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
di: O'Reilly, Cliff, et al.
Pubblicazione: (2025)
di: O'Reilly, Cliff, et al.
Pubblicazione: (2025)
Experiments in News Bias Detection with Pre-Trained Neural Transformers
di: Menzner, Tim, et al.
Pubblicazione: (2024)
di: Menzner, Tim, et al.
Pubblicazione: (2024)
Unifying Structured Data as Graph for Data-to-Text Pre-Training
di: Li, Shujie, et al.
Pubblicazione: (2024)
di: Li, Shujie, et al.
Pubblicazione: (2024)
Rethinking Reflection in Pre-Training
di: AI, Essential, et al.
Pubblicazione: (2025)
di: AI, Essential, et al.
Pubblicazione: (2025)
Crown, Frame, Reverse: Layer-Wise Scaling Variants for LLM Pre-Training
di: Baroian, Andrei, et al.
Pubblicazione: (2025)
di: Baroian, Andrei, et al.
Pubblicazione: (2025)
Towards an automatic method for generating topical vocabulary test forms for specific reading passages
di: Flor, Michael, et al.
Pubblicazione: (2025)
di: Flor, Michael, et al.
Pubblicazione: (2025)
Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
di: Jiang, Changhao, et al.
Pubblicazione: (2025)
di: Jiang, Changhao, et al.
Pubblicazione: (2025)
Reinforcement Learning on Pre-Training Data
di: Li, Siheng, et al.
Pubblicazione: (2025)
di: Li, Siheng, et al.
Pubblicazione: (2025)
Continual Pre-Training for Cross-Lingual LLM Adaptation: Enhancing Japanese Language Capabilities
di: Fujii, Kazuki, et al.
Pubblicazione: (2024)
di: Fujii, Kazuki, et al.
Pubblicazione: (2024)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
di: Du, Wenyu, et al.
Pubblicazione: (2024)
di: Du, Wenyu, et al.
Pubblicazione: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
di: Weller, Orion, et al.
Pubblicazione: (2023)
di: Weller, Orion, et al.
Pubblicazione: (2023)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
di: Ceron, Tanise, et al.
Pubblicazione: (2025)
di: Ceron, Tanise, et al.
Pubblicazione: (2025)
Investigating Gender Bias in LLM-Generated Stories via Psychological Stereotypes
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
di: Masoudian, Shahed, et al.
Pubblicazione: (2025)
Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance
di: Zheng, Weihua, et al.
Pubblicazione: (2026)
di: Zheng, Weihua, et al.
Pubblicazione: (2026)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
Unraveling Emotions with Pre-Trained Models
di: Pajón-Sanmartín, Alejandro, et al.
Pubblicazione: (2025)
di: Pajón-Sanmartín, Alejandro, et al.
Pubblicazione: (2025)
SlimPajama-DC: Understanding Data Combinations for LLM Training
di: Shen, Zhiqiang, et al.
Pubblicazione: (2023)
di: Shen, Zhiqiang, et al.
Pubblicazione: (2023)
Large Language Model Capabilities in Perioperative Risk Prediction and Prognostication
di: Chung, Philip, et al.
Pubblicazione: (2024)
di: Chung, Philip, et al.
Pubblicazione: (2024)
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models
di: Tian, Junfeng, et al.
Pubblicazione: (2024)
di: Tian, Junfeng, et al.
Pubblicazione: (2024)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
di: Zhuo, Le, et al.
Pubblicazione: (2024)
di: Zhuo, Le, et al.
Pubblicazione: (2024)
Evolution of Concepts in Language Model Pre-Training
di: Ge, Xuyang, et al.
Pubblicazione: (2025)
di: Ge, Xuyang, et al.
Pubblicazione: (2025)
Beyond Memorization: The Challenge of Random Memory Access in Language Models
di: Zhu, Tongyao, et al.
Pubblicazione: (2024)
di: Zhu, Tongyao, et al.
Pubblicazione: (2024)
Scaling LLM Pre-training with Vocabulary Curriculum
di: Yu, Fangyuan
Pubblicazione: (2025)
di: Yu, Fangyuan
Pubblicazione: (2025)
D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training
di: Xu, Yuanjian, et al.
Pubblicazione: (2026)
di: Xu, Yuanjian, et al.
Pubblicazione: (2026)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
di: Chung, Yi-Ling, et al.
Pubblicazione: (2025)
DEFT: Data Efficient Fine-Tuning for Pre-Trained Language Models via Unsupervised Core-Set Selection
di: Das, Devleena, et al.
Pubblicazione: (2023)
di: Das, Devleena, et al.
Pubblicazione: (2023)
Revealing the Inherent Instructability of Pre-Trained Language Models
di: An, Seokhyun, et al.
Pubblicazione: (2024)
di: An, Seokhyun, et al.
Pubblicazione: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis
di: Matta, Shiho, et al.
Pubblicazione: (2024)
di: Matta, Shiho, et al.
Pubblicazione: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
di: Aynetdinov, Ansar, et al.
Pubblicazione: (2025)
di: Aynetdinov, Ansar, et al.
Pubblicazione: (2025)
Pre-Trained Language Models for Keyphrase Prediction: A Review
di: Umair, Muhammad, et al.
Pubblicazione: (2024)
di: Umair, Muhammad, et al.
Pubblicazione: (2024)
Characterizing Stereotypical Bias from Privacy-preserving Pre-Training
di: Arnold, Stefan, et al.
Pubblicazione: (2024)
di: Arnold, Stefan, et al.
Pubblicazione: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
di: Qiu, Zeju, et al.
Pubblicazione: (2025)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
di: Kim, Eunsu, et al.
Pubblicazione: (2024)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
di: Fleshman, William, et al.
Pubblicazione: (2024)
di: Fleshman, William, et al.
Pubblicazione: (2024)
MTQE.en-he: Machine Translation Quality Estimation for English-Hebrew
di: Rosenbaum, Andy, et al.
Pubblicazione: (2026)
di: Rosenbaum, Andy, et al.
Pubblicazione: (2026)
Fixing It in Post: A Comparative Study of LLM Post-Training Data Quality and Model Performance
di: Djuhera, Aladin, et al.
Pubblicazione: (2025)
di: Djuhera, Aladin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Attribution Crisis in LLM Search Results
di: Strauss, Ilan, et al.
Pubblicazione: (2025) -
Real-World Gaps in AI Governance Research
di: Strauss, Ilan, et al.
Pubblicazione: (2025) -
LLM-Supported Natural Language to Bash Translation
di: Westenfelder, Finnian, et al.
Pubblicazione: (2025) -
Disentangling concept semantics via multilingual averaging in Sparse Autoencoders
di: O'Reilly, Cliff, et al.
Pubblicazione: (2025) -
Experiments in News Bias Detection with Pre-Trained Neural Transformers
di: Menzner, Tim, et al.
Pubblicazione: (2024)