Montessori-Instruct: Generate Influential Training Data Tailored for Student Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaochuan, Yu, Zichun, Xiong, Chenyan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026)
by: Yu, Zichun, et al.
Published: (2026)
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024)
by: Yu, Zichun, et al.
Published: (2024)
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025)
by: Wang, Tevin, et al.
Published: (2025)
PithTrain: A Compact and Agent-Native MoE Training System
by: Lai, Ruihang, et al.
Published: (2026)
by: Lai, Ruihang, et al.
Published: (2026)
Influential Language Data Selection via Gradient Trajectory Pursuit
by: Deng, Zhiwei, et al.
Published: (2024)
by: Deng, Zhiwei, et al.
Published: (2024)
FLAME-MoE: A Transparent End-to-End Research Platform for Mixture-of-Experts Language Models
by: Kang, Hao, et al.
Published: (2025)
by: Kang, Hao, et al.
Published: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
IPCGRL: Language-Instructed Reinforcement Learning for Procedural Level Generation
by: Baek, In-Chang, et al.
Published: (2025)
by: Baek, In-Chang, et al.
Published: (2025)
ED-Copilot: Reduce Emergency Department Wait Time with Language Model Diagnostic Assistance
by: Sun, Liwen, et al.
Published: (2024)
by: Sun, Liwen, et al.
Published: (2024)
Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning
by: Yu, Erxin, et al.
Published: (2025)
by: Yu, Erxin, et al.
Published: (2025)
AgentInstruct: Toward Generative Teaching with Agentic Flows
by: Mitra, Arindam, et al.
Published: (2024)
by: Mitra, Arindam, et al.
Published: (2024)
Agent Instructs Large Language Models to be General Zero-Shot Reasoners
by: Crispino, Nicholas, et al.
Published: (2023)
by: Crispino, Nicholas, et al.
Published: (2023)
Ultra-Fast Language Generation via Discrete Diffusion Divergence Instruct
by: Zheng, Haoyang, et al.
Published: (2025)
by: Zheng, Haoyang, et al.
Published: (2025)
Reinforcement Learning on Pre-Training Data
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
Beyond Input Activations: Identifying Influential Latents by Gradient Sparse Autoencoders
by: Shu, Dong, et al.
Published: (2025)
by: Shu, Dong, et al.
Published: (2025)
CodecLM: Aligning Language Models with Tailored Synthetic Data
by: Wang, Zifeng, et al.
Published: (2024)
by: Wang, Zifeng, et al.
Published: (2024)
RAG-Instruct: Boosting LLMs with Diverse Retrieval-Augmented Instructions
by: Liu, Wanlong, et al.
Published: (2024)
by: Liu, Wanlong, et al.
Published: (2024)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
by: Toshniwal, Shubham, et al.
Published: (2024)
by: Toshniwal, Shubham, et al.
Published: (2024)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Inspection and Control of Self-Generated-Text Recognition Ability in Llama3-8b-Instruct
by: Ackerman, Christopher, et al.
Published: (2024)
by: Ackerman, Christopher, et al.
Published: (2024)
Group-Level Data Selection for Efficient Pretraining
by: Yu, Zichun, et al.
Published: (2025)
by: Yu, Zichun, et al.
Published: (2025)
Domain Adaptation of Llama3-70B-Instruct through Continual Pre-Training and Model Merging: A Comprehensive Evaluation
by: Siriwardhana, Shamane, et al.
Published: (2024)
by: Siriwardhana, Shamane, et al.
Published: (2024)
A Regularization-based Transfer Learning Method for Information Extraction via Instructed Graph Decoder
by: Chen, Kedi, et al.
Published: (2024)
by: Chen, Kedi, et al.
Published: (2024)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Text2Data: Low-Resource Data Generation with Textual Control
by: Wang, Shiyu, et al.
Published: (2024)
by: Wang, Shiyu, et al.
Published: (2024)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
by: Khan, Zaid, et al.
Published: (2024)
by: Khan, Zaid, et al.
Published: (2024)
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
by: Zhang, Jianguo, et al.
Published: (2024)
by: Zhang, Jianguo, et al.
Published: (2024)
InstructAV: Instruction Fine-tuning Large Language Models for Authorship Verification
by: Hu, Yujia, et al.
Published: (2024)
by: Hu, Yujia, et al.
Published: (2024)
Instruct-Tuning Pretrained Causal Language Models for Ancient Greek Papyrology and Epigraphy
by: Cullhed, Eric
Published: (2024)
by: Cullhed, Eric
Published: (2024)
Scalability Matters: Overcoming Challenges in InstructGLM with Similarity-Degree-Based Sampling
by: Lee, Hyun, et al.
Published: (2025)
by: Lee, Hyun, et al.
Published: (2025)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
by: Tang, Zihao, et al.
Published: (2024)
by: Tang, Zihao, et al.
Published: (2024)
Training and Evaluating Language Models with Template-based Data Generation
by: Zhang, Yifan
Published: (2024)
by: Zhang, Yifan
Published: (2024)
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs
by: Ali, Mehdi, et al.
Published: (2024)
by: Ali, Mehdi, et al.
Published: (2024)
LAM SIMULATOR: Advancing Data Generation for Large Action Model Training via Online Exploration and Trajectory Feedback
by: Hoang, Thai, et al.
Published: (2025)
by: Hoang, Thai, et al.
Published: (2025)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
by: Toshniwal, Shubham, et al.
Published: (2024)
by: Toshniwal, Shubham, et al.
Published: (2024)
Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque
by: Corral, Ander, et al.
Published: (2024)
by: Corral, Ander, et al.
Published: (2024)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
by: Luo, Haipeng, et al.
Published: (2023)
by: Luo, Haipeng, et al.
Published: (2023)
Learning to Detect Language Model Training Data via Active Reconstruction
by: Yin, Junjie Oscar, et al.
Published: (2026)
by: Yin, Junjie Oscar, et al.
Published: (2026)
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
Similar Items
-
Generating Pretraining Tokens from Organic Data for Data-Bound Scaling
by: Yu, Zichun, et al.
Published: (2026) -
RePro: Training Language Models to Faithfully Recycle the Web for Pretraining
by: Yu, Zichun, et al.
Published: (2025) -
MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence Models
by: Yu, Zichun, et al.
Published: (2024) -
AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning
by: Wang, Tevin, et al.
Published: (2025) -
PithTrain: A Compact and Agent-Native MoE Training System
by: Lai, Ruihang, et al.
Published: (2026)