Skill-Targeted Adaptive Training
Fuente:
arXiv
Saved in:
| Main Authors: | He, Yinghui, Panigrahi, Abhishek, Lin, Yong, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025)
by: He, Yinghui, et al.
Published: (2025)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022)
by: Malladi, Sadhika, et al.
Published: (2022)
Why is Your Language Model a Poor Implicit Reward Model?
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
On the Power of Context-Enhanced Learning in LLMs
by: Zhu, Xingyu, et al.
Published: (2025)
by: Zhu, Xingyu, et al.
Published: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
by: Xia, Mengzhou, et al.
Published: (2024)
by: Xia, Mengzhou, et al.
Published: (2024)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
by: Didolkar, Aniket, et al.
Published: (2025)
by: Didolkar, Aniket, et al.
Published: (2025)
Provable unlearning in topic modeling and downstream tasks
by: Wei, Stanley, et al.
Published: (2024)
by: Wei, Stanley, et al.
Published: (2024)
Representing Rule-based Chatbots with Transformers
by: Friedman, Dan, et al.
Published: (2024)
by: Friedman, Dan, et al.
Published: (2024)
AI-Assisted Generation of Difficult Math Questions
by: Shah, Vedant, et al.
Published: (2024)
by: Shah, Vedant, et al.
Published: (2024)
When Errors Can Be Beneficial: A Categorization of Imperfect Rewards for Policy Gradient
by: Shang, Shuning, et al.
Published: (2026)
by: Shang, Shuning, et al.
Published: (2026)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
by: Zhao, Haoyu, et al.
Published: (2025)
by: Zhao, Haoyu, et al.
Published: (2025)
Trainable Transformer in Transformer
by: Panigrahi, Abhishek, et al.
Published: (2023)
by: Panigrahi, Abhishek, et al.
Published: (2023)
Goedel-Prover: A Frontier Model for Open-Source Automated Theorem Proving
by: Lin, Yong, et al.
Published: (2025)
by: Lin, Yong, et al.
Published: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
by: Cheng, Yun, et al.
Published: (2026)
by: Cheng, Yun, et al.
Published: (2026)
Unrealized Expectations: Comparing AI Methods vs Classical Algorithms for Maximum Independent Set
by: Wu, Yikai, et al.
Published: (2025)
by: Wu, Yikai, et al.
Published: (2025)
Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity
by: Costales, Robby, et al.
Published: (2024)
by: Costales, Robby, et al.
Published: (2024)
ReSkill: Reconciling Skill Creation with Policy Optimization in Agentic RL
by: He, Zelin, et al.
Published: (2026)
by: He, Zelin, et al.
Published: (2026)
Building Better Deception Probes Using Targeted Instruction Pairs
by: Natarajan, Vikram, et al.
Published: (2026)
by: Natarajan, Vikram, et al.
Published: (2026)
SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training
by: He, Zhongyu, et al.
Published: (2026)
by: He, Zhongyu, et al.
Published: (2026)
LiBaGS: Lightweight Boundary Gap Synthesis for Targeted Synthetic Data Selection
by: Moturu, Abhishek, et al.
Published: (2026)
by: Moturu, Abhishek, et al.
Published: (2026)
REE-TTT: Highly Adaptive Radar Echo Extrapolation Based on Test-Time Training
by: Di, Xin, et al.
Published: (2026)
by: Di, Xin, et al.
Published: (2026)
Unlearning via Sparse Representations
by: Shah, Vedant, et al.
Published: (2023)
by: Shah, Vedant, et al.
Published: (2023)
Adaptive Action Duration with Contextual Bandits for Deep Reinforcement Learning in Dynamic Environments
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
APAR: Modeling Irregular Target Functions in Tabular Regression via Arithmetic-Aware Pre-Training and Adaptive-Regularized Fine-Tuning
by: Wu, Hong-Wei, et al.
Published: (2024)
by: Wu, Hong-Wei, et al.
Published: (2024)
Generalized Adaptive Transfer Network: Enhancing Transfer Learning in Reinforcement Learning Across Domains
by: Verma, Abhishek, et al.
Published: (2025)
by: Verma, Abhishek, et al.
Published: (2025)
Optimistic Verifiable Training by Controlling Hardware Nondeterminism
by: Srivastava, Megha, et al.
Published: (2024)
by: Srivastava, Megha, et al.
Published: (2024)
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
by: Manivannan, Sanjeev, et al.
Published: (2026)
by: Manivannan, Sanjeev, et al.
Published: (2026)
Self-Adaptive Graph Mixture of Models
by: Meena, Mohit, et al.
Published: (2025)
by: Meena, Mohit, et al.
Published: (2025)
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
by: Lyu, Kaifeng, et al.
Published: (2024)
by: Lyu, Kaifeng, et al.
Published: (2024)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
On the Impossibility of Retrain Equivalence in Machine Unlearning
by: Yu, Jiatong, et al.
Published: (2025)
by: Yu, Jiatong, et al.
Published: (2025)
Goal Exploration via Adaptive Skill Distribution for Goal-Conditioned Reinforcement Learning
by: Wu, Lisheng, et al.
Published: (2024)
by: Wu, Lisheng, et al.
Published: (2024)
Highly Parallelized Reinforcement Learning Training with Relaxed Assignment Dependencies
by: He, Zhouyu, et al.
Published: (2025)
by: He, Zhouyu, et al.
Published: (2025)
MelissaDL x Breed: Towards Data-Efficient On-line Supervised Training of Multi-parametric Surrogates with Active Learning
by: Dymchenko, Sofya, et al.
Published: (2024)
by: Dymchenko, Sofya, et al.
Published: (2024)
Stable On-Policy Distillation through Adaptive Target Reformulation
by: Jang, Ijun, et al.
Published: (2026)
by: Jang, Ijun, et al.
Published: (2026)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
Rethinking Thinking Tokens: LLMs as Improvement Operators
by: Madaan, Lovish, et al.
Published: (2025)
by: Madaan, Lovish, et al.
Published: (2025)
Act as You Learn: Adaptive Decision-Making in Non-Stationary Markov Decision Processes
by: Luo, Baiting, et al.
Published: (2024)
by: Luo, Baiting, et al.
Published: (2024)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
by: Razin, Noam, et al.
Published: (2025)
by: Razin, Noam, et al.
Published: (2025)
Similar Items
-
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
by: He, Yinghui, et al.
Published: (2025) -
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
by: Malladi, Sadhika, et al.
Published: (2022) -
Why is Your Language Model a Poor Implicit Reward Model?
by: Razin, Noam, et al.
Published: (2025) -
Can Models Learn Skill Composition from Examples?
by: Zhao, Haoyu, et al.
Published: (2024) -
On the Power of Context-Enhanced Learning in LLMs
by: Zhu, Xingyu, et al.
Published: (2025)