Rethinking the Role of Text Complexity in Language Model Pretraining
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Velasco, Dan John, Roque, Matthew Theodore |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
von: Roque, Matthew Theodore, et al.
Veröffentlicht: (2025)
von: Roque, Matthew Theodore, et al.
Veröffentlicht: (2025)
Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
von: Velasco, Dan John, et al.
Veröffentlicht: (2025)
von: Velasco, Dan John, et al.
Veröffentlicht: (2025)
Harnessing the Intrinsic Knowledge of Pretrained Language Models for Challenging Text Classification Settings
von: Gao, Lingyu
Veröffentlicht: (2024)
von: Gao, Lingyu
Veröffentlicht: (2024)
Drop Dropout on Single-Epoch Language Model Pretraining
von: Liu, Houjun, et al.
Veröffentlicht: (2025)
von: Liu, Houjun, et al.
Veröffentlicht: (2025)
Multilingual Pretraining for Pixel Language Models
von: Kesen, Ilker, et al.
Veröffentlicht: (2025)
von: Kesen, Ilker, et al.
Veröffentlicht: (2025)
Rethinking the Role of Proxy Rewards in Language Model Alignment
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
Generalizable and Stable Finetuning of Pretrained Language Models on Low-Resource Texts
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
von: Somayajula, Sai Ashish, et al.
Veröffentlicht: (2024)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
von: Verma, Vivek, et al.
Veröffentlicht: (2023)
von: Verma, Vivek, et al.
Veröffentlicht: (2023)
Parrot Mind: Towards Explaining the Complex Task Reasoning of Pretrained Large Language Models with Template-Content Structure
von: Yang, Haotong, et al.
Veröffentlicht: (2023)
von: Yang, Haotong, et al.
Veröffentlicht: (2023)
ConSiDERS-The-Human Evaluation Framework: Rethinking Human Evaluation for Generative Large Language Models
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
von: Elangovan, Aparna, et al.
Veröffentlicht: (2024)
On Linear Representations and Pretraining Data Frequency in Language Models
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
von: Merullo, Jack, et al.
Veröffentlicht: (2025)
Gender Encoding Patterns in Pretrained Language Model Representations
von: Zakizadeh, Mahdi, et al.
Veröffentlicht: (2025)
von: Zakizadeh, Mahdi, et al.
Veröffentlicht: (2025)
PonderLM: Pretraining Language Models to Ponder in Continuous Space
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
von: Zeng, Boyi, et al.
Veröffentlicht: (2025)
Code Pretraining Improves Entity Tracking Abilities of Language Models
von: Kim, Najoung, et al.
Veröffentlicht: (2024)
von: Kim, Najoung, et al.
Veröffentlicht: (2024)
Emergent Abilities of Large Language Models under Continued Pretraining for Language Adaptation
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
von: Elhady, Ahmed, et al.
Veröffentlicht: (2025)
BioPars: A Pretrained Biomedical Large Language Model for Persian Biomedical Text Mining
von: Merzah, Baqer M., et al.
Veröffentlicht: (2025)
von: Merzah, Baqer M., et al.
Veröffentlicht: (2025)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
von: Lee, Ivan, et al.
Veröffentlicht: (2025)
Large Language Models for Biomedical Text Simplification: Promising But Not There Yet
von: Li, Zihao, et al.
Veröffentlicht: (2024)
von: Li, Zihao, et al.
Veröffentlicht: (2024)
UrbanCLIP: Learning Text-enhanced Urban Region Profiling with Contrastive Language-Image Pretraining from the Web
von: Yan, Yibo, et al.
Veröffentlicht: (2023)
von: Yan, Yibo, et al.
Veröffentlicht: (2023)
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning
von: Yu, Huimu, et al.
Veröffentlicht: (2024)
von: Yu, Huimu, et al.
Veröffentlicht: (2024)
A Causal Language Modeling Detour Improves Encoder Continued Pretraining
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
von: Touchent, Rian, et al.
Veröffentlicht: (2026)
Exploring the Benefits of Domain-Pretraining of Generative Large Language Models for Chemistry
von: Acharya, Anurag, et al.
Veröffentlicht: (2024)
von: Acharya, Anurag, et al.
Veröffentlicht: (2024)
Generative Pretrained Structured Transformers: Unsupervised Syntactic Language Models at Scale
von: Hu, Xiang, et al.
Veröffentlicht: (2024)
von: Hu, Xiang, et al.
Veröffentlicht: (2024)
Rethinking Text-based Protein Understanding: Retrieval or LLM?
von: Wu, Juntong, et al.
Veröffentlicht: (2025)
von: Wu, Juntong, et al.
Veröffentlicht: (2025)
Rethinking Word Similarity: Semantic Similarity through Classification Confusion
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2025)
Rethinking DPO: The Role of Rejected Responses in Preference Misalignment
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Jay Hyeon, et al.
Veröffentlicht: (2025)
Exploring Polyglot Harmony: On Multilingual Data Allocation for Large Language Models Pretraining
von: Guo, Ping, et al.
Veröffentlicht: (2025)
von: Guo, Ping, et al.
Veröffentlicht: (2025)
Memory Decoder: A Pretrained, Plug-and-Play Memory for Large Language Models
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
von: Cao, Jiaqi, et al.
Veröffentlicht: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
von: Zhang, Yang, et al.
Veröffentlicht: (2025)
Systematic Analysis for Pretrained Language Model Priming for Parameter-Efficient Fine-tuning
von: Huang, Shih-Cheng, et al.
Veröffentlicht: (2022)
von: Huang, Shih-Cheng, et al.
Veröffentlicht: (2022)
Foundation CAN LM: A Pretrained Language Model For Automotive CAN Data
von: Esashi, Akiharu, et al.
Veröffentlicht: (2026)
von: Esashi, Akiharu, et al.
Veröffentlicht: (2026)
AraPoemBERT: A Pretrained Language Model for Arabic Poetry Analysis
von: Qarah, Faisal
Veröffentlicht: (2024)
von: Qarah, Faisal
Veröffentlicht: (2024)
Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
von: Du, Xinrun, et al.
Veröffentlicht: (2024)
von: Du, Xinrun, et al.
Veröffentlicht: (2024)
SaudiBERT: A Large Language Model Pretrained on Saudi Dialect Corpora
von: Qarah, Faisal
Veröffentlicht: (2024)
von: Qarah, Faisal
Veröffentlicht: (2024)
Deciphering the Impact of Pretraining Data on Large Language Models through Machine Unlearning
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
von: Zhao, Yang, et al.
Veröffentlicht: (2024)
Rethinking the Outlier Distribution in Large Language Models: An In-depth Study
von: Raman, Rahul, et al.
Veröffentlicht: (2025)
von: Raman, Rahul, et al.
Veröffentlicht: (2025)
Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
von: McLeish, Sean, et al.
Veröffentlicht: (2025)
Towards Text-free Graph Foundation Models: Rethinking Multi-Domain Graph Contrastive Learning
von: Zhao, Zihao, et al.
Veröffentlicht: (2025)
von: Zhao, Zihao, et al.
Veröffentlicht: (2025)
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment
von: Agarwalla, Abhinav, et al.
Veröffentlicht: (2024)
von: Agarwalla, Abhinav, et al.
Veröffentlicht: (2024)
Dynamic Masking Rate Schedules for MLM Pretraining
von: Ankner, Zachary, et al.
Veröffentlicht: (2023)
von: Ankner, Zachary, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
von: Roque, Matthew Theodore, et al.
Veröffentlicht: (2025) -
Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
von: Velasco, Dan John, et al.
Veröffentlicht: (2025) -
Harnessing the Intrinsic Knowledge of Pretrained Language Models for Challenging Text Classification Settings
von: Gao, Lingyu
Veröffentlicht: (2024) -
Drop Dropout on Single-Epoch Language Model Pretraining
von: Liu, Houjun, et al.
Veröffentlicht: (2025) -
Multilingual Pretraining for Pixel Language Models
von: Kesen, Ilker, et al.
Veröffentlicht: (2025)