SkillFactory: Self-Distillation For Learning Cognitive Behaviors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sprague, Zayne, Lu, Jack, Wadhwa, Manya, Keh, Sedrick, Ren, Mengye, Durrett, Greg |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025)
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
von: Sprague, Zayne, et al.
Veröffentlicht: (2024)
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2023)
von: Sprague, Zayne, et al.
Veröffentlicht: (2023)
Learning to Refine with Fine-Grained Natural Language Feedback
von: Wadhwa, Manya, et al.
Veröffentlicht: (2024)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2024)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024)
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024)
Using Natural Language Explanations to Rescale Human Judgments
von: Wadhwa, Manya, et al.
Veröffentlicht: (2023)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2023)
SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation
von: Divekar, Abhishek, et al.
Veröffentlicht: (2024)
von: Divekar, Abhishek, et al.
Veröffentlicht: (2024)
Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents
von: Ding, Wenxuan, et al.
Veröffentlicht: (2026)
von: Ding, Wenxuan, et al.
Veröffentlicht: (2026)
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
von: Tang, Liyan, et al.
Veröffentlicht: (2024)
Context Tuning for In-Context Optimization
von: Lu, Jack, et al.
Veröffentlicht: (2025)
von: Lu, Jack, et al.
Veröffentlicht: (2025)
A Critical Evaluation of AI Feedback for Aligning Large Language Models
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
von: Sharma, Archit, et al.
Veröffentlicht: (2024)
Learning Composable Chains-of-Thought
von: Yin, Fangcong, et al.
Veröffentlicht: (2025)
von: Yin, Fangcong, et al.
Veröffentlicht: (2025)
Contrastive Learning to Improve Retrieval for Real-world Fact Checking
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
von: Sriram, Aniruddh, et al.
Veröffentlicht: (2024)
CREATE: Testing LLMs for Associative Creativity
von: Wadhwa, Manya, et al.
Veröffentlicht: (2026)
von: Wadhwa, Manya, et al.
Veröffentlicht: (2026)
PropMEND: Hypernetworks for Knowledge Propagation in LLMs
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
von: Liu, Zeyu Leo, et al.
Veröffentlicht: (2025)
CoLLEGe: Concept Embedding Generation for Large Language Models
von: Teehan, Ryan, et al.
Veröffentlicht: (2024)
von: Teehan, Ryan, et al.
Veröffentlicht: (2024)
Adaptive Margin RLHF via Preference over Preferences
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
von: Chittepu, Yaswanth, et al.
Veröffentlicht: (2025)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
von: Huang, Jiazhen, et al.
Veröffentlicht: (2026)
QUDsim: Quantifying Discourse Similarities in LLM-Generated Text
von: Namuduri, Ramya, et al.
Veröffentlicht: (2025)
von: Namuduri, Ramya, et al.
Veröffentlicht: (2025)
Distilling Event Sequence Knowledge From Large Language Models
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
von: Wadhwa, Somin, et al.
Veröffentlicht: (2024)
Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
von: Dai, Hui, et al.
Veröffentlicht: (2024)
von: Dai, Hui, et al.
Veröffentlicht: (2024)
Is the Top Still Spinning? Evaluating Subjectivity in Narrative Understanding
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
von: Subbiah, Melanie, et al.
Veröffentlicht: (2025)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
von: Wang, Hao, et al.
Veröffentlicht: (2026)
von: Wang, Hao, et al.
Veröffentlicht: (2026)
Memory Transfer Learning: How Memories are Transferred Across Domains in Coding Agents
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
SkillOS: Learning Skill Curation for Self-Evolving Agents
von: Ouyang, Siru, et al.
Veröffentlicht: (2026)
von: Ouyang, Siru, et al.
Veröffentlicht: (2026)
MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
von: Che, Xinyu, et al.
Veröffentlicht: (2026)
ProofWala: A Framework for Multilingual Proof Data Synthesis and Theorem-Proving
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
Where It Really Matters: Few-Shot Environmental Conservation Media Monitoring for Low-Resource Languages
von: Jain, Sameer, et al.
Veröffentlicht: (2024)
von: Jain, Sameer, et al.
Veröffentlicht: (2024)
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
von: Tang, Liyan, et al.
Veröffentlicht: (2025)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Self-Knowledge Distillation for Learning Ambiguity
von: Park, Hancheol, et al.
Veröffentlicht: (2024)
von: Park, Hancheol, et al.
Veröffentlicht: (2024)
Aligning LLMs with Human Uncertainty: A Beta-Bernoulli Calibrator for LLM Forecasting
von: Dai, Hui, et al.
Veröffentlicht: (2026)
von: Dai, Hui, et al.
Veröffentlicht: (2026)
TSUBASA: Improving Long-Horizon Personalization via Evolving Memory and Self-Learning with Context Distillation
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2026)
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2026)
EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial Negotiation
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
von: Long, Yunbo, et al.
Veröffentlicht: (2026)
OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
von: Hallinan, Skyler, et al.
Veröffentlicht: (2026)
von: Hallinan, Skyler, et al.
Veröffentlicht: (2026)
SkillOpt: Executive Strategy for Self-Evolving Agent Skills
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
von: Yang, Yifan, et al.
Veröffentlicht: (2026)
TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM
von: Zhou, Haoyang, et al.
Veröffentlicht: (2026)
von: Zhou, Haoyang, et al.
Veröffentlicht: (2026)
ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
von: Zhu, Jie, et al.
Veröffentlicht: (2026)
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
von: Zhang, Haozhen, et al.
Veröffentlicht: (2026)
von: Zhang, Haozhen, et al.
Veröffentlicht: (2026)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2024)
von: Liu, Biao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EvalAgent: Discovering Implicit Evaluation Criteria from the Web
von: Wadhwa, Manya, et al.
Veröffentlicht: (2025) -
To CoT or not to CoT? Chain-of-thought helps mainly on math and symbolic reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2024) -
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
von: Sprague, Zayne, et al.
Veröffentlicht: (2023) -
Learning to Refine with Fine-Grained Natural Language Feedback
von: Wadhwa, Manya, et al.
Veröffentlicht: (2024) -
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
von: Gunjal, Anisha, et al.
Veröffentlicht: (2024)