Can Models Learn Skill Composition from Examples?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Haoyu, Kaur, Simran, Yu, Dingli, Goyal, Anirudh, Arora, Sanjeev |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
von: Lyu, Kaifeng, et al.
Veröffentlicht: (2024)
von: Lyu, Kaifeng, et al.
Veröffentlicht: (2024)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
von: Kaur, Simran, et al.
Veröffentlicht: (2024)
von: Kaur, Simran, et al.
Veröffentlicht: (2024)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)
Why is Your Language Model a Poor Implicit Reward Model?
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
von: Didolkar, Aniket, et al.
Veröffentlicht: (2025)
von: Didolkar, Aniket, et al.
Veröffentlicht: (2025)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
von: Park, Simon, et al.
Veröffentlicht: (2025)
von: Park, Simon, et al.
Veröffentlicht: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
von: Guo, Siyuan, et al.
Veröffentlicht: (2024)
von: Guo, Siyuan, et al.
Veröffentlicht: (2024)
The Role of Diversity in In-Context Learning for Large Language Models
von: Xiao, Wenyang, et al.
Veröffentlicht: (2025)
von: Xiao, Wenyang, et al.
Veröffentlicht: (2025)
LoLCATs: On Low-Rank Linearizing of Large Language Models
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
von: Zhang, Michael, et al.
Veröffentlicht: (2024)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
von: Razin, Noam, et al.
Veröffentlicht: (2025)
von: Razin, Noam, et al.
Veröffentlicht: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
von: Xia, Mengzhou, et al.
Veröffentlicht: (2024)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
von: Razin, Noam, et al.
Veröffentlicht: (2024)
von: Razin, Noam, et al.
Veröffentlicht: (2024)
AI-Assisted Generation of Difficult Math Questions
von: Shah, Vedant, et al.
Veröffentlicht: (2024)
von: Shah, Vedant, et al.
Veröffentlicht: (2024)
LoRA Users Beware: A Few Spurious Tokens Can Manipulate Your Finetuned Model
von: Salles, Marcel Mateos, et al.
Veröffentlicht: (2025)
von: Salles, Marcel Mateos, et al.
Veröffentlicht: (2025)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
von: Yao, Xinhao, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Reasoning in Large Language Models with One Training Example
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
von: Wang, Yiping, et al.
Veröffentlicht: (2025)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
von: Goyal, Naman
Veröffentlicht: (2024)
von: Goyal, Naman
Veröffentlicht: (2024)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
von: Potamitis, Nearchos, et al.
Veröffentlicht: (2025)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
von: Zhang, Haozhen, et al.
Veröffentlicht: (2026)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2026)
RetICL: Sequential Retrieval of In-Context Examples with Reinforcement Learning
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2023)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2023)
Scaling Test-Time Compute for Agentic Coding
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
von: Kim, Joongwon, et al.
Veröffentlicht: (2026)
Debate Helps Weak Judges Reward Stronger Models
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
Unfamiliar Finetuning Examples Control How Language Models Hallucinate
von: Kang, Katie, et al.
Veröffentlicht: (2024)
von: Kang, Katie, et al.
Veröffentlicht: (2024)
Characterizing Model-Native Skills
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
von: Kang, Feiyang, et al.
Veröffentlicht: (2026)
Cartridges: Lightweight and general-purpose long context representations via self-study
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
von: Eyuboglu, Sabri, et al.
Veröffentlicht: (2025)
Principled Data Selection for Alignment: The Hidden Risks of Difficult Examples
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
von: Gao, Chengqian, et al.
Veröffentlicht: (2025)
Learn while Unlearn: An Iterative Unlearning Framework for Generative Language Models
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
von: Tang, Haoyu, et al.
Veröffentlicht: (2024)
Self-Play with Adversarial Critic: Provable and Scalable Offline Alignment for Language Models
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
von: Ji, Xiang, et al.
Veröffentlicht: (2024)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Can Large Language Models Unlock Novel Scientific Research Ideas?
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
von: Kumar, Sandeep, et al.
Veröffentlicht: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
von: Zhang, Jiajie, et al.
Veröffentlicht: (2025)
ContextFocus: Activation Steering for Contextual Faithfulness in Large Language Models
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
von: Anand, Nikhil, et al.
Veröffentlicht: (2026)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2025)
Skill-Targeted Adaptive Training
von: He, Yinghui, et al.
Veröffentlicht: (2025)
von: He, Yinghui, et al.
Veröffentlicht: (2025)
Memento-Skills: Let Agents Design Agents
von: Zhou, Huichi, et al.
Veröffentlicht: (2026)
von: Zhou, Huichi, et al.
Veröffentlicht: (2026)
Learning to Route for Dynamic Adapter Composition in Continual Learning with Language Models
von: Araujo, Vladimir, et al.
Veröffentlicht: (2024)
von: Araujo, Vladimir, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
von: Lyu, Kaifeng, et al.
Veröffentlicht: (2024) -
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
von: Kaur, Simran, et al.
Veröffentlicht: (2024) -
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
von: Park, Simon, et al.
Veröffentlicht: (2025) -
Contextual Drag: How Errors in the Context Affect LLM Reasoning
von: Cheng, Yun, et al.
Veröffentlicht: (2026) -
Ineq-Comp: Benchmarking Human-Intuitive Compositional Reasoning in Automated Theorem Proving on Inequalities
von: Zhao, Haoyu, et al.
Veröffentlicht: (2025)