CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Butt, Natasha, Manczak, Blazej, Wiggers, Auke, Rainone, Corrado, Zhang, David W., Defferrard, Michaël, Cohen, Taco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
von: Ding, Liang
Veröffentlicht: (2026)
von: Ding, Liang
Veröffentlicht: (2026)
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
von: Manczak, Blazej, et al.
Veröffentlicht: (2024)
von: Manczak, Blazej, et al.
Veröffentlicht: (2024)
Prioritized Replay for RL Post-training
von: Fatemi, Mehdi
Veröffentlicht: (2026)
von: Fatemi, Mehdi
Veröffentlicht: (2026)
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
von: Shianifar, Jonaid, et al.
Veröffentlicht: (2026)
von: Shianifar, Jonaid, et al.
Veröffentlicht: (2026)
Replacing thinking with tool usage enables reasoning in small language models
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
Shallow Robustness, Deep Vulnerabilities: Multi-Turn Evaluation of Medical LLMs
von: Manczak, Blazej, et al.
Veröffentlicht: (2025)
von: Manczak, Blazej, et al.
Veröffentlicht: (2025)
Contact Energy Based Hindsight Experience Prioritization
von: Sayar, Erdi, et al.
Veröffentlicht: (2023)
von: Sayar, Erdi, et al.
Veröffentlicht: (2023)
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
von: Yeo, Woongyeng, et al.
Veröffentlicht: (2026)
SD-Search: On-Policy Hindsight Self-Distillation for Search-Augmented Reasoning
von: Ma, Yufei, et al.
Veröffentlicht: (2026)
von: Ma, Yufei, et al.
Veröffentlicht: (2026)
Extrapolative Weight Averaging Reveals Correctness-Efficiency Frontiers in Code RL
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2026)
What Makes Large Language Models Reason in (Multi-Turn) Code Generation?
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
von: Zheng, Kunhao, et al.
Veröffentlicht: (2024)
When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
von: Rainone, Corrado, et al.
Veröffentlicht: (2026)
von: Rainone, Corrado, et al.
Veröffentlicht: (2026)
Provable Interactive Learning with Hindsight Instruction Feedback
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
von: Misra, Dipendra, et al.
Veröffentlicht: (2024)
RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
Contextual Experience Replay for Self-Improvement of Language Agents
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
von: Liu, Yitao, et al.
Veröffentlicht: (2025)
Enabling Option Learning in Sparse Rewards with Hindsight Experience Replay
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
von: Romio, Gabriel, et al.
Veröffentlicht: (2026)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
von: Hu, Michael Y., et al.
Veröffentlicht: (2025)
Adaptable Hindsight Experience Replay for Search-Based Learning
von: Vazaios, Alexandros, et al.
Veröffentlicht: (2025)
von: Vazaios, Alexandros, et al.
Veröffentlicht: (2025)
Code Pretraining Improves Entity Tracking Abilities of Language Models
von: Kim, Najoung, et al.
Veröffentlicht: (2024)
von: Kim, Najoung, et al.
Veröffentlicht: (2024)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Improving Language Model Reasoning with Self-motivated Learning
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
von: Feng, Yunlong, et al.
Veröffentlicht: (2024)
Improving Retrieval Augmented Language Model with Self-Reasoning
von: Xia, Yuan, et al.
Veröffentlicht: (2024)
von: Xia, Yuan, et al.
Veröffentlicht: (2024)
ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay
von: Hu, Zhexin, et al.
Veröffentlicht: (2026)
von: Hu, Zhexin, et al.
Veröffentlicht: (2026)
Self-Improving World Modelling with Latent Actions
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
Self-generated Replay Memories for Continual Neural Machine Translation
von: Resta, Michele, et al.
Veröffentlicht: (2024)
von: Resta, Michele, et al.
Veröffentlicht: (2024)
On the "Induction Bias" in Sequence Models
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2026)
von: Ebrahimi, M. Reza, et al.
Veröffentlicht: (2026)
Soft Tokens, Hard Truths
von: Butt, Natasha, et al.
Veröffentlicht: (2025)
von: Butt, Natasha, et al.
Veröffentlicht: (2025)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness
von: Ivanov, Igor, et al.
Veröffentlicht: (2026)
von: Ivanov, Igor, et al.
Veröffentlicht: (2026)
FOREVER: Forgetting Curve-Inspired Memory Replay for Language Model Continual Learning
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
von: Feng, Yujie, et al.
Veröffentlicht: (2026)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
von: Abbes, Istabrak, et al.
Veröffentlicht: (2025)
Cluster-based Sampling in Hindsight Experience Replay for Robotic Tasks (Student Abstract)
von: Kim, Taeyoung, et al.
Veröffentlicht: (2022)
von: Kim, Taeyoung, et al.
Veröffentlicht: (2022)
Investigating the Interplay of Prioritized Replay and Generalization
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
von: Panahi, Parham Mohammad, et al.
Veröffentlicht: (2024)
Self-Execution Simulation Improves Coding Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
Advancing Large Language Model Attribution through Self-Improving
von: Huang, Lei, et al.
Veröffentlicht: (2024)
von: Huang, Lei, et al.
Veröffentlicht: (2024)
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
von: Arnesen, Samuel, et al.
Veröffentlicht: (2024)
von: Arnesen, Samuel, et al.
Veröffentlicht: (2024)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
von: Sancaktar, Cansu, et al.
Veröffentlicht: (2026)
von: Sancaktar, Cansu, et al.
Veröffentlicht: (2026)
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyin, et al.
Veröffentlicht: (2024)
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding
von: Zhao, Xuanle, et al.
Veröffentlicht: (2026)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AgentHER: Hindsight Experience Replay for LLM Agent Trajectory Relabeling
von: Ding, Liang
Veröffentlicht: (2026) -
PrimeGuard: Safe and Helpful LLMs through Tuning-Free Routing
von: Manczak, Blazej, et al.
Veröffentlicht: (2024) -
Prioritized Replay for RL Post-training
von: Fatemi, Mehdi
Veröffentlicht: (2026) -
Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning
von: Shianifar, Jonaid, et al.
Veröffentlicht: (2026) -
Replacing thinking with tool usage enables reasoning in small language models
von: Rainone, Corrado, et al.
Veröffentlicht: (2025)