Synthetic Sandbox for Training Machine Learning Engineering Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Yuhang, Zhang, Lizhu, Wu, Yifan, Liu, Jiayi, Fan, Xiangjun, Zhao, Zhuokai, Yan, Hong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
by: Zhou, Yuhang, et al.
Published: (2026)
by: Zhou, Yuhang, et al.
Published: (2026)
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
by: Zeng, Hanqing, et al.
Published: (2025)
by: Zeng, Hanqing, et al.
Published: (2025)
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026)
by: Rai, Arushi, et al.
Published: (2026)
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
by: Yang, Chenghao, et al.
Published: (2025)
by: Yang, Chenghao, et al.
Published: (2025)
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)
by: Xiong, Nuoya, et al.
Published: (2026)
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
by: Han, Kevin, et al.
Published: (2026)
by: Han, Kevin, et al.
Published: (2026)
Agentic Recommender System with Hierarchical Belief-State Memory
by: Shen, Xiang, et al.
Published: (2026)
by: Shen, Xiang, et al.
Published: (2026)
GEM: Empowering LLM for both Embedding Generation and Language Understanding
by: Zhang, Caojin, et al.
Published: (2025)
by: Zhang, Caojin, et al.
Published: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
by: Ruan, Yangjun, et al.
Published: (2023)
by: Ruan, Yangjun, et al.
Published: (2023)
Mixture-of-Minds: Multi-Agent Reinforcement Learning for Table Understanding
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
ML-Agent: Reinforcing LLM Agents for Autonomous Machine Learning Engineering
by: Liu, Zexi, et al.
Published: (2025)
by: Liu, Zexi, et al.
Published: (2025)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
by: Yu, Hao, et al.
Published: (2025)
by: Yu, Hao, et al.
Published: (2025)
AutoEnv: Automated Environments for Measuring Cross-Environment Agent Learning
by: Zhang, Jiayi, et al.
Published: (2025)
by: Zhang, Jiayi, et al.
Published: (2025)
SYN-DIGITS: A Synthetic Control Framework for Calibrated Digital Twin Simulation
by: Fan, Grace Jiarui, et al.
Published: (2026)
by: Fan, Grace Jiarui, et al.
Published: (2026)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
by: Zhou, Yifei, et al.
Published: (2024)
by: Zhou, Yifei, et al.
Published: (2024)
HyperAdaLoRA: Accelerating LoRA Rank Allocation During Training via Hypernetworks without Sacrificing Performance
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning
by: Golubev, Alexander, et al.
Published: (2025)
by: Golubev, Alexander, et al.
Published: (2025)
WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning
by: Wei, Zhepei, et al.
Published: (2025)
by: Wei, Zhepei, et al.
Published: (2025)
Learning Virtual Machine Scheduling in Cloud Computing through Language Agents
by: Wu, JieHao, et al.
Published: (2025)
by: Wu, JieHao, et al.
Published: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
by: Zhao, Shiwan, et al.
Published: (2026)
by: Zhao, Shiwan, et al.
Published: (2026)
SynthAgent: Adapting Web Agents with Synthetic Supervision
by: Wang, Zhaoyang, et al.
Published: (2025)
by: Wang, Zhaoyang, et al.
Published: (2025)
Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe
by: Wu, Xixi, et al.
Published: (2026)
by: Wu, Xixi, et al.
Published: (2026)
Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch
by: Zhang, Ziyang, et al.
Published: (2025)
by: Zhang, Ziyang, et al.
Published: (2025)
Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
Training and Evaluating Language Models with Template-based Data Generation
by: Zhang, Yifan
Published: (2024)
by: Zhang, Yifan
Published: (2024)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
by: Xu, Ran, et al.
Published: (2026)
by: Xu, Ran, et al.
Published: (2026)
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
by: Li, Zongxia, et al.
Published: (2025)
by: Li, Zongxia, et al.
Published: (2025)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
by: Wang, Chaoqi, et al.
Published: (2025)
by: Wang, Chaoqi, et al.
Published: (2025)
Vocal Sandbox: Continual Learning and Adaptation for Situated Human-Robot Collaboration
by: Grannen, Jennifer, et al.
Published: (2024)
by: Grannen, Jennifer, et al.
Published: (2024)
SELA: Tree-Search Enhanced LLM Agents for Automated Machine Learning
by: Chi, Yizhou, et al.
Published: (2024)
by: Chi, Yizhou, et al.
Published: (2024)
Retrieval-Reasoning Large Language Model-based Synthetic Clinical Trial Generation
by: Xu, Zerui, et al.
Published: (2024)
by: Xu, Zerui, et al.
Published: (2024)
Sarcasm Detection on Reddit Using Classical Machine Learning and Feature Engineering
by: Karmaker, Subrata
Published: (2025)
by: Karmaker, Subrata
Published: (2025)
Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning
by: Wang, Zhaoyang, et al.
Published: (2026)
by: Wang, Zhaoyang, et al.
Published: (2026)
Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
by: Ma, Wenhan, et al.
Published: (2025)
by: Ma, Wenhan, et al.
Published: (2025)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
AgentOhana: Design Unified Data and Training Pipeline for Effective Agent Learning
by: Zhang, Jianguo, et al.
Published: (2024)
by: Zhang, Jianguo, et al.
Published: (2024)
FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers
by: Wang, Sihan, et al.
Published: (2026)
by: Wang, Sihan, et al.
Published: (2026)
ReMix: Reinforcement routing for mixtures of LoRAs in LLM finetuning
by: Qiu, Ruizhong, et al.
Published: (2026)
by: Qiu, Ruizhong, et al.
Published: (2026)
Similar Items
-
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
by: Zhou, Yuhang, et al.
Published: (2026) -
S'MoRE: Structural Mixture of Residual Experts for Parameter-Efficient LLM Fine-tuning
by: Zeng, Hanqing, et al.
Published: (2025) -
TARo: Token-level Adaptive Routing for LLM Test-time Alignment
by: Rai, Arushi, et al.
Published: (2026) -
Let it Calm: Exploratory Annealed Decoding for Verifiable Reinforcement Learning
by: Yang, Chenghao, et al.
Published: (2025) -
Token-Level LLM Collaboration via FusionRoute
by: Xiong, Nuoya, et al.
Published: (2026)