Evolution of Concepts in Language Model Pre-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Ge, Xuyang, Shu, Wentao, Wu, Jiaxing, Zhou, Yunhua, He, Zhengfu, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024)
by: Wang, Junxuan, et al.
Published: (2024)
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025)
by: Wang, Junxuan, et al.
Published: (2025)
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024)
by: Ge, Xuyang, et al.
Published: (2024)
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
by: Zhou, Guancheng, et al.
Published: (2026)
by: Zhou, Guancheng, et al.
Published: (2026)
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)
by: Ye, Jiasheng, et al.
Published: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
Explicit Multi-head Attention for Inter-head Interaction in Large Language Models
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
BitStack: Any-Size Compression of Large Language Models in Variable Memory Environments
by: Wang, Xinghao, et al.
Published: (2024)
by: Wang, Xinghao, et al.
Published: (2024)
Balanced Data Sampling for Language Model Training with Clustering
by: Shao, Yunfan, et al.
Published: (2024)
by: Shao, Yunfan, et al.
Published: (2024)
Towards Understanding the Nature of Attention with Low-Rank Sparse Decomposition
by: He, Zhengfu, et al.
Published: (2025)
by: He, Zhengfu, et al.
Published: (2025)
Evaluating the Performance of Large Language Models on GAOKAO Benchmark
by: Zhang, Xiaotian, et al.
Published: (2023)
by: Zhang, Xiaotian, et al.
Published: (2023)
Can AI Assistants Know What They Don't Know?
by: Cheng, Qinyuan, et al.
Published: (2024)
by: Cheng, Qinyuan, et al.
Published: (2024)
Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
by: He, Zhengfu, et al.
Published: (2024)
by: He, Zhengfu, et al.
Published: (2024)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
by: Xu, Ningyu, et al.
Published: (2026)
by: Xu, Ningyu, et al.
Published: (2026)
Improving Large Language Models with Concept-Aware Fine-Tuning
by: Chen, Michael K., et al.
Published: (2025)
by: Chen, Michael K., et al.
Published: (2025)
GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation
by: Zong, Yi, et al.
Published: (2024)
by: Zong, Yi, et al.
Published: (2024)
Mousse: Rectifying the Geometry of Muon with Curvature-Aware Preconditioning
by: Zhang, Yechen, et al.
Published: (2026)
by: Zhang, Yechen, et al.
Published: (2026)
ConceptMath: A Bilingual Concept-wise Benchmark for Measuring Mathematical Reasoning of Large Language Models
by: Wu, Yanan, et al.
Published: (2024)
by: Wu, Yanan, et al.
Published: (2024)
Revealing the Inherent Instructability of Pre-Trained Language Models
by: An, Seokhyun, et al.
Published: (2024)
by: An, Seokhyun, et al.
Published: (2024)
Topic Over Source: The Key to Effective Data Mixing for Language Models Pre-training
by: Peng, Jiahui, et al.
Published: (2025)
by: Peng, Jiahui, et al.
Published: (2025)
Domain2Vec: Vectorizing Datasets to Find the Optimal Data Mixture without Training
by: Zhang, Mozhi, et al.
Published: (2025)
by: Zhang, Mozhi, et al.
Published: (2025)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
by: Aynetdinov, Ansar, et al.
Published: (2025)
by: Aynetdinov, Ansar, et al.
Published: (2025)
Pre-Trained Language Models for Keyphrase Prediction: A Review
by: Umair, Muhammad, et al.
Published: (2024)
by: Umair, Muhammad, et al.
Published: (2024)
Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration
by: Sun, Qiushi, et al.
Published: (2023)
by: Sun, Qiushi, et al.
Published: (2023)
ProtLLM: An Interleaved Protein-Language LLM with Protein-as-Word Pre-Training
by: Zhuo, Le, et al.
Published: (2024)
by: Zhuo, Le, et al.
Published: (2024)
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
by: Liu, Yuliang, et al.
Published: (2026)
by: Liu, Yuliang, et al.
Published: (2026)
Interactive Evolution: A Neural-Symbolic Self-Training Framework For Large Language Models
by: Xu, Fangzhi, et al.
Published: (2024)
by: Xu, Fangzhi, et al.
Published: (2024)
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics
by: Wei, Ting-Ruen, et al.
Published: (2025)
by: Wei, Ting-Ruen, et al.
Published: (2025)
How to Set the Learning Rate for Large-Scale Pre-training?
by: Zhou, Yunhua, et al.
Published: (2026)
by: Zhou, Yunhua, et al.
Published: (2026)
Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
by: Huber, Bernd, et al.
Published: (2025)
by: Huber, Bernd, et al.
Published: (2025)
Legal Documents Drafting with Fine-Tuned Pre-Trained Large Language Model
by: Lin, Chun-Hsien, et al.
Published: (2024)
by: Lin, Chun-Hsien, et al.
Published: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
by: Guigon, Maxime, et al.
Published: (2026)
by: Guigon, Maxime, et al.
Published: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
Turn Waste into Worth: Rectifying Top-$k$ Router of MoE
by: Zeng, Zhiyuan, et al.
Published: (2024)
by: Zeng, Zhiyuan, et al.
Published: (2024)
Recursive Concept Evolution for Compositional Reasoning in Large Language Models
by: Chaudhry, Sarim
Published: (2026)
by: Chaudhry, Sarim
Published: (2026)
Unraveling Emotions with Pre-Trained Models
by: Pajón-Sanmartín, Alejandro, et al.
Published: (2025)
by: Pajón-Sanmartín, Alejandro, et al.
Published: (2025)
Agent Alignment in Evolving Social Norms
by: Li, Shimin, et al.
Published: (2024)
by: Li, Shimin, et al.
Published: (2024)
Safe Inputs but Unsafe Output: Benchmarking Cross-modality Safety Alignment of Large Vision-Language Model
by: Wang, Siyin, et al.
Published: (2024)
by: Wang, Siyin, et al.
Published: (2024)
Construction of Hyper-Relational Knowledge Graphs Using Pre-Trained Large Language Models
by: Datta, Preetha, et al.
Published: (2024)
by: Datta, Preetha, et al.
Published: (2024)
Similar Items
-
Towards Universality: Studying Mechanistic Similarity Across Language Model Architectures
by: Wang, Junxuan, et al.
Published: (2024) -
Dimensional Collapse in Transformer Attention Outputs: A Challenge for Sparse Dictionary Learning
by: Wang, Junxuan, et al.
Published: (2025) -
Automatically Identifying Local and Global Circuits with Linear Computation Graphs
by: Ge, Xuyang, et al.
Published: (2024) -
A Distributional View for Visual Mechanistic Interpretability: KL-Minimal Soft-Constraint Principle
by: Zhou, Guancheng, et al.
Published: (2026) -
Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance
by: Ye, Jiasheng, et al.
Published: (2024)