Sub-goal Distillation: A Method to Improve Small Language Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Hashemzadeh, Maryam, Stengel-Eskin, Elias, Chandar, Sarath, Cote, Marc-Alexandre |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026)
Soft Self-Consistency Improves Language Model Agents
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Language-guided Skill Learning with Temporal Variational Inference
di: Fu, Haotian, et al.
Pubblicazione: (2024)
di: Fu, Haotian, et al.
Pubblicazione: (2024)
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024)
di: Wan, David, et al.
Pubblicazione: (2024)
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
di: Prasad, Archiki, et al.
Pubblicazione: (2023)
DataEnvGym: Data Generation Agents in Teacher Environments with Student Feedback
di: Khan, Zaid, et al.
Pubblicazione: (2024)
di: Khan, Zaid, et al.
Pubblicazione: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
UPCORE: Utility-Preserving Coreset Selection for Balanced Unlearning
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
di: Patil, Vaidehi, et al.
Pubblicazione: (2025)
Multi-Attribute Steering of Language Models via Targeted Intervention
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
di: Nguyen, Duy, et al.
Pubblicazione: (2025)
Faithfulness Measurable Masked Language Models
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
di: Madsen, Andreas, et al.
Pubblicazione: (2023)
Are self-explanations from Large Language Models faithful?
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
di: Madsen, Andreas, et al.
Pubblicazione: (2024)
ReGAL: Refactoring Programs to Discover Generalizable Abstractions
di: Stengel-Eskin, Elias, et al.
Pubblicazione: (2024)
di: Stengel-Eskin, Elias, et al.
Pubblicazione: (2024)
LASeR: Learning to Adaptively Select Reward Models with Multi-Armed Bandits
di: Nguyen, Duy, et al.
Pubblicazione: (2024)
di: Nguyen, Duy, et al.
Pubblicazione: (2024)
Exploring Quantization for Efficient Pre-Training of Transformer Language Models
di: Chitsaz, Kamran, et al.
Pubblicazione: (2024)
di: Chitsaz, Kamran, et al.
Pubblicazione: (2024)
Probabilistic Calibration Is a Trainable Capability in Language Models
di: Baldelli, Davide, et al.
Pubblicazione: (2026)
di: Baldelli, Davide, et al.
Pubblicazione: (2026)
GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization
di: Khan, Zaid, et al.
Pubblicazione: (2026)
di: Khan, Zaid, et al.
Pubblicazione: (2026)
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
di: Xiao, Hanqi, et al.
Pubblicazione: (2026)
di: Xiao, Hanqi, et al.
Pubblicazione: (2026)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
CoPeP: Benchmarking Continual Pretraining for Protein Language Models
di: Patil, Darshan, et al.
Pubblicazione: (2026)
di: Patil, Darshan, et al.
Pubblicazione: (2026)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
The Expressive Limits of Diagonal SSMs for State-Tracking
di: Shakerinava, Mehran, et al.
Pubblicazione: (2026)
di: Shakerinava, Mehran, et al.
Pubblicazione: (2026)
NovoMolGen: Rethinking Molecular Language Model Pretraining
di: Chitsaz, Kamran, et al.
Pubblicazione: (2025)
di: Chitsaz, Kamran, et al.
Pubblicazione: (2025)
Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
One Life to Learn: Inferring Symbolic World Models for Stochastic Environments from Unguided Exploration
di: Khan, Zaid, et al.
Pubblicazione: (2025)
di: Khan, Zaid, et al.
Pubblicazione: (2025)
Shielded Controller Units for RL with Operational Constraints Applied to Remote Microgrids
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
di: Nekoei, Hadi, et al.
Pubblicazione: (2025)
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
di: Chen, Justin Chih-Yao, et al.
Pubblicazione: (2024)
Mastering Memory Tasks with World Models
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
di: Samsami, Mohammad Reza, et al.
Pubblicazione: (2024)
Manifold Metric: A Loss Landscape Approach for Predicting Model Performance
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
di: Malviya, Pranshu, et al.
Pubblicazione: (2024)
Steering Large Language Model Activations in Sparse Spaces
di: Bayat, Reza, et al.
Pubblicazione: (2025)
di: Bayat, Reza, et al.
Pubblicazione: (2025)
Intelligent Switching for Reset-Free RL
di: Patil, Darshan, et al.
Pubblicazione: (2024)
di: Patil, Darshan, et al.
Pubblicazione: (2024)
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
di: Guiroy, Simon, et al.
Pubblicazione: (2025)
Lookbehind-SAM: k steps back, 1 step forward
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
di: Mordido, Gonçalo, et al.
Pubblicazione: (2023)
Conflict-Resolving and Sharpness-Aware Minimization for Generalized Knowledge Editing with Multiple Updates
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
System-1.x: Learning to Balance Fast and Slow Planning with Language Models
di: Saha, Swarnadeep, et al.
Pubblicazione: (2024)
di: Saha, Swarnadeep, et al.
Pubblicazione: (2024)
Do Large Language Models Know How Much They Know?
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
di: Prato, Gabriele, et al.
Pubblicazione: (2025)
Reconstruction or Semantics? What Makes a Latent Space Useful for Robotic World Models
di: Nilaksh, et al.
Pubblicazione: (2026)
di: Nilaksh, et al.
Pubblicazione: (2026)
BindGPT: A Scalable Framework for 3D Molecular Design via Language Modeling and Reinforcement Learning
di: Zholus, Artem, et al.
Pubblicazione: (2024)
di: Zholus, Artem, et al.
Pubblicazione: (2024)
Towards Practical Tool Usage for Continually Learning LLMs
di: Huang, Jerry, et al.
Pubblicazione: (2024)
di: Huang, Jerry, et al.
Pubblicazione: (2024)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
di: Nilaksh, et al.
Pubblicazione: (2026)
di: Nilaksh, et al.
Pubblicazione: (2026)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
di: Nguyen, Duy, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing
di: Hashemzadeh, Maryam, et al.
Pubblicazione: (2026) -
Soft Self-Consistency Improves Language Model Agents
di: Wang, Han, et al.
Pubblicazione: (2024) -
Language-guided Skill Learning with Temporal Variational Inference
di: Fu, Haotian, et al.
Pubblicazione: (2024) -
Contrastive Region Guidance: Improving Grounding in Vision-Language Models without Training
di: Wan, David, et al.
Pubblicazione: (2024) -
Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language Models
di: Prasad, Archiki, et al.
Pubblicazione: (2023)