HRM-Text: Efficient Pretraining Beyond Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Guan, Liu, Changling, Wang, Chenyu, Zhou, Cai, Sun, Yuhao, Wu, Yifei, Zhen, Shuai, Scimeca, Luca, Yadkori, Yasin Abbasi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Why Attend to Everything? Focus is the Key
by: Yao, Hengshuai, et al.
Published: (2026)
by: Yao, Hengshuai, et al.
Published: (2026)
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025)
by: Wang, Guan, et al.
Published: (2025)
Pointwise confidence estimation in the non-linear $\ell^2$-regularized least squares
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
Low-rank bias, weight decay, and model merging in neural networks
by: Kuzborskij, Ilja, et al.
Published: (2025)
by: Kuzborskij, Ilja, et al.
Published: (2025)
Efficient Pretraining Length Scaling
by: Wu, Bohong, et al.
Published: (2025)
by: Wu, Bohong, et al.
Published: (2025)
Next Semantic Scale Prediction via Hierarchical Diffusion Language Models
by: Zhou, Cai, et al.
Published: (2025)
by: Zhou, Cai, et al.
Published: (2025)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
by: Huang, Han, et al.
Published: (2024)
by: Huang, Han, et al.
Published: (2024)
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
by: Wang, Yifei
Published: (2025)
by: Wang, Yifei
Published: (2025)
Beyond Repetition: Text Simplification and Curriculum Learning for Data-Constrained Pretraining
by: Roque, Matthew Theodore, et al.
Published: (2025)
by: Roque, Matthew Theodore, et al.
Published: (2025)
Best of both worlds: Stochastic & adversarial best-arm identification
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
by: Abbasi-Yadkori, Yasin, et al.
Published: (2026)
Efficient Symbolic Execution of Software under Fault Attacks
by: Fang, Yuzhou, et al.
Published: (2025)
by: Fang, Yuzhou, et al.
Published: (2025)
BeyondWeb: Lessons from Scaling Synthetic Data for Trillion-scale Pretraining
by: DatologyAI, et al.
Published: (2025)
by: DatologyAI, et al.
Published: (2025)
QuaDMix: Quality-Diversity Balanced Data Selection for Efficient LLM Pretraining
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
Scaling, Simplification, and Adaptation: Lessons from Pretraining on Machine-Translated Text
by: Velasco, Dan John, et al.
Published: (2025)
by: Velasco, Dan John, et al.
Published: (2025)
A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization
by: Sun, Dan, et al.
Published: (2025)
by: Sun, Dan, et al.
Published: (2025)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
by: Zheng, Hao, et al.
Published: (2025)
by: Zheng, Hao, et al.
Published: (2025)
Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
by: Bai, Tianyi, et al.
Published: (2024)
by: Bai, Tianyi, et al.
Published: (2024)
Look Ahead or Look Around? A Theoretical Comparison Between Autoregressive and Masked Pretraining
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
by: Li, Jeffrey, et al.
Published: (2026)
by: Li, Jeffrey, et al.
Published: (2026)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Mitigating LLM Hallucinations via Conformal Abstention
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
by: Yadkori, Yasin Abbasi, et al.
Published: (2024)
Beyond Text Compression: Evaluating Tokenizers Across Scales
by: Lotz, Jonas F., et al.
Published: (2025)
by: Lotz, Jonas F., et al.
Published: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Beyond Length: Quantifying Long-Range Information for Long-Context LLM Pretraining Data
by: Deng, Haoran, et al.
Published: (2025)
by: Deng, Haoran, et al.
Published: (2025)
Beyond Random Sampling: Efficient Language Model Pretraining via Curriculum Learning
by: Zhang, Yang, et al.
Published: (2025)
by: Zhang, Yang, et al.
Published: (2025)
FRAME: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
MathPile: A Billion-Token-Scale Pretraining Corpus for Math
by: Wang, Zengzhi, et al.
Published: (2023)
by: Wang, Zengzhi, et al.
Published: (2023)
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
by: Cen, Zhepeng, et al.
Published: (2025)
by: Cen, Zhepeng, et al.
Published: (2025)
WRAP++: Web discoveRy Amplified Pretraining
by: Zhou, Jiang, et al.
Published: (2026)
by: Zhou, Jiang, et al.
Published: (2026)
Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers
by: Gong, Linyuan, et al.
Published: (2023)
by: Gong, Linyuan, et al.
Published: (2023)
AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse Decoding
by: Luo, Shuqing, et al.
Published: (2025)
by: Luo, Shuqing, et al.
Published: (2025)
Preference Curriculum: LLMs Should Always Be Pretrained on Their Preferred Data
by: Zhang, Xuemiao, et al.
Published: (2025)
by: Zhang, Xuemiao, et al.
Published: (2025)
Table-to-Text Generation with Pretrained Diffusion Models
by: Krylov, Aleksei S., et al.
Published: (2024)
by: Krylov, Aleksei S., et al.
Published: (2024)
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026)
by: Miok, Kristian, et al.
Published: (2026)
Pushing The Limit of LLM Capacity for Text Classification
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Target-Oriented Pretraining Data Selection via Neuron-Activated Graph
by: Wang, Zijun, et al.
Published: (2026)
by: Wang, Zijun, et al.
Published: (2026)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
by: Fang, Rongyao, et al.
Published: (2025)
by: Fang, Rongyao, et al.
Published: (2025)
PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining
by: Zhu, Guanghao, et al.
Published: (2026)
by: Zhu, Guanghao, et al.
Published: (2026)
Similar Items
-
Why Attend to Everything? Focus is the Key
by: Yao, Hengshuai, et al.
Published: (2026) -
To Believe or Not to Believe Your LLM
by: Yadkori, Yasin Abbasi, et al.
Published: (2024) -
Hierarchical Reasoning Model
by: Wang, Guan, et al.
Published: (2025) -
Pointwise confidence estimation in the non-linear $\ell^2$-regularized least squares
by: Kuzborskij, Ilja, et al.
Published: (2025) -
Low-rank bias, weight decay, and model merging in neural networks
by: Kuzborskij, Ilja, et al.
Published: (2025)