Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zelikman, Eric, Lorch, Eliana, Mackey, Lester, Kalai, Adam Tauman |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Consensus Sampling for Safer Generative AI
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2025)
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2025)
Do Language Models Know When They're Hallucinating References?
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023)
Calibrated Language Models Must Hallucinate
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
von: Suzgun, Mirac, et al.
Veröffentlicht: (2024)
von: Suzgun, Mirac, et al.
Veröffentlicht: (2024)
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)
UnStar: Unlearning with Self-Taught Anti-Sample Reasoning for LLMs
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
von: Sinha, Yash, et al.
Veröffentlicht: (2024)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
von: Koh, Woosung, et al.
Veröffentlicht: (2025)
von: Koh, Woosung, et al.
Veröffentlicht: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2024)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
von: Sahoo, Subramanyam, et al.
Veröffentlicht: (2026)
On Non-interactive Evaluation of Animal Communication Translators
von: Paradise, Orr, et al.
Veröffentlicht: (2025)
von: Paradise, Orr, et al.
Veröffentlicht: (2025)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
von: Xiong, Feng, et al.
Veröffentlicht: (2025)
von: Xiong, Feng, et al.
Veröffentlicht: (2025)
Hypothesis Search: Inductive Reasoning with Language Models
von: Wang, Ruocheng, et al.
Veröffentlicht: (2023)
von: Wang, Ruocheng, et al.
Veröffentlicht: (2023)
Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking
von: Zelikman, Eric, et al.
Veröffentlicht: (2024)
von: Zelikman, Eric, et al.
Veröffentlicht: (2024)
CodeIt: Self-Improving Language Models with Prioritized Hindsight Replay
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
von: Butt, Natasha, et al.
Veröffentlicht: (2024)
Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
von: Rosset, Corby, et al.
Veröffentlicht: (2024)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
von: Wen, Xueru, et al.
Veröffentlicht: (2025)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
Self-Taught Recognizer: Toward Unsupervised Adaptation for Speech Foundation Models
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
von: Hu, Yuchen, et al.
Veröffentlicht: (2024)
West-of-N: Synthetic Preferences for Self-Improving Reward Models
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
von: Pace, Alizée, et al.
Veröffentlicht: (2024)
Self-Taught Evaluators
von: Wang, Tianlu, et al.
Veröffentlicht: (2024)
von: Wang, Tianlu, et al.
Veröffentlicht: (2024)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
Adapting Language Models via Token Translation
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
von: Feng, Zhili, et al.
Veröffentlicht: (2024)
Large Language Models Must Be Taught to Know What They Don't Know
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
von: Kapoor, Sanyam, et al.
Veröffentlicht: (2024)
Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long Context
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2026)
Self-Consistency Preference Optimization
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
von: Prasad, Archiki, et al.
Veröffentlicht: (2024)
Self-Supervised Prompt Optimization
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
von: Xiang, Jinyu, et al.
Veröffentlicht: (2025)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2024)
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2024)
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents
von: Li, Hanchen, et al.
Veröffentlicht: (2026)
von: Li, Hanchen, et al.
Veröffentlicht: (2026)
Self-Improving LLM Agents at Test-Time
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
Self-Improving World Modelling with Latent Actions
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
von: Liang, Yu, et al.
Veröffentlicht: (2026)
von: Liang, Yu, et al.
Veröffentlicht: (2026)
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
von: Akter, Syeda Nahida, et al.
Veröffentlicht: (2024)
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
von: Gupta, Sonam, et al.
Veröffentlicht: (2024)
Hexa: Self-Improving for Knowledge-Grounded Dialogue System
von: Jo, Daejin, et al.
Veröffentlicht: (2023)
von: Jo, Daejin, et al.
Veröffentlicht: (2023)
Soft Self-Consistency Improves Language Model Agents
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
TSO: Self-Training with Scaled Preference Optimization
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
von: Chen, Kaihui, et al.
Veröffentlicht: (2024)
Recursive Agent Optimization
von: Gandhi, Apurva, et al.
Veröffentlicht: (2026)
von: Gandhi, Apurva, et al.
Veröffentlicht: (2026)
Reasoning Distillation and Structural Alignment for Improved Code Generation
von: Jalilifard, Amir, et al.
Veröffentlicht: (2025)
von: Jalilifard, Amir, et al.
Veröffentlicht: (2025)
Integrative Decoding: Improve Factuality via Implicit Self-consistency
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
von: Cheng, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Consensus Sampling for Safer Generative AI
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2025) -
Do Language Models Know When They're Hallucinating References?
von: Agrawal, Ayush, et al.
Veröffentlicht: (2023) -
Calibrated Language Models Must Hallucinate
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023) -
Meta-Prompting: Enhancing Language Models with Task-Agnostic Scaffolding
von: Suzgun, Mirac, et al.
Veröffentlicht: (2024) -
V-STaR: Training Verifiers for Self-Taught Reasoners
von: Hosseini, Arian, et al.
Veröffentlicht: (2024)