ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Kim, Joongwon, Goyal, Anirudh, Tan, Liang, Hajishirzi, Hannaneh, Iyer, Srinivasan, Wang, Tianlu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
par: Kim, Joongwon, et autres
Publié: (2024)
par: Kim, Joongwon, et autres
Publié: (2024)
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
par: Kim, Joongwon, et autres
Publié: (2024)
par: Kim, Joongwon, et autres
Publié: (2024)
Data Engineering for Scaling Language Models to 128K Context
par: Fu, Yao, et autres
Publié: (2024)
par: Fu, Yao, et autres
Publié: (2024)
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
par: Lyu, Xinxi, et autres
Publié: (2024)
par: Lyu, Xinxi, et autres
Publié: (2024)
BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
par: Cao, Qingqing, et autres
Publié: (2023)
par: Cao, Qingqing, et autres
Publié: (2023)
Scaling Test-Time Compute for Agentic Coding
par: Kim, Joongwon, et autres
Publié: (2026)
par: Kim, Joongwon, et autres
Publié: (2026)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
par: Graf, Victoria, et autres
Publié: (2026)
par: Graf, Victoria, et autres
Publié: (2026)
OLMES: A Standard for Language Model Evaluations
par: Gu, Yuling, et autres
Publié: (2024)
par: Gu, Yuling, et autres
Publié: (2024)
Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
par: Liu, Jiacheng, et autres
Publié: (2024)
par: Liu, Jiacheng, et autres
Publié: (2024)
Emergent Search and Backtracking in Latent Reasoning Models
par: Cui, Jasmine, et autres
Publié: (2026)
par: Cui, Jasmine, et autres
Publié: (2026)
Step Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
par: Yang, Xiao-Wen, et autres
Publié: (2025)
par: Yang, Xiao-Wen, et autres
Publié: (2025)
Learning to Detect Language Model Training Data via Active Reconstruction
par: Yin, Junjie Oscar, et autres
Publié: (2026)
par: Yin, Junjie Oscar, et autres
Publié: (2026)
OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization
par: Sun, Yiyou, et autres
Publié: (2025)
par: Sun, Yiyou, et autres
Publié: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
par: Wiegreffe, Sarah, et autres
Publié: (2024)
par: Wiegreffe, Sarah, et autres
Publié: (2024)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
par: Lu, Pan, et autres
Publié: (2023)
par: Lu, Pan, et autres
Publié: (2023)
Steering When Necessary: Flexible Steering Large Language Models with Backtracking
par: Cheng, Zifeng, et autres
Publié: (2025)
par: Cheng, Zifeng, et autres
Publié: (2025)
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
par: Min, Sewon, et autres
Publié: (2023)
par: Min, Sewon, et autres
Publié: (2023)
Reliable, Adaptable, and Attributable Language Models with Retrieval
par: Asai, Akari, et autres
Publié: (2024)
par: Asai, Akari, et autres
Publié: (2024)
Backtracking for Safety
par: Sel, Bilgehan, et autres
Publié: (2025)
par: Sel, Bilgehan, et autres
Publié: (2025)
BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
par: Wu, Qinzhuo, et autres
Publié: (2025)
par: Wu, Qinzhuo, et autres
Publié: (2025)
Teaching Language Models to Reason with Tools
par: Li, Chengpeng, et autres
Publié: (2025)
par: Li, Chengpeng, et autres
Publié: (2025)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
par: Wang, Bing, et autres
Publié: (2026)
par: Wang, Bing, et autres
Publié: (2026)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
par: Chen, Tong, et autres
Publié: (2025)
par: Chen, Tong, et autres
Publié: (2025)
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
par: Wang, Jingyuan, et autres
Publié: (2025)
par: Wang, Jingyuan, et autres
Publié: (2025)
Fluid Language Model Benchmarking
par: Hofmann, Valentin, et autres
Publié: (2025)
par: Hofmann, Valentin, et autres
Publié: (2025)
Bayesian Teaching Enables Probabilistic Reasoning in Large Language Models
par: Qiu, Linlu, et autres
Publié: (2025)
par: Qiu, Linlu, et autres
Publié: (2025)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
par: Liu, Jiacheng, et autres
Publié: (2023)
par: Liu, Jiacheng, et autres
Publié: (2023)
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models
par: Tan, Wenting, et autres
Publié: (2024)
par: Tan, Wenting, et autres
Publié: (2024)
TasTe: Teaching Large Language Models to Translate through Self-Reflection
par: Wang, Yutong, et autres
Publié: (2024)
par: Wang, Yutong, et autres
Publié: (2024)
Correction with Backtracking Reduces Hallucination in Summarization
par: Liu, Zhenzhen, et autres
Publié: (2023)
par: Liu, Zhenzhen, et autres
Publié: (2023)
Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge
par: Saha, Swarnadeep, et autres
Publié: (2025)
par: Saha, Swarnadeep, et autres
Publié: (2025)
ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
par: Tanmay, Kumar, et autres
Publié: (2025)
par: Tanmay, Kumar, et autres
Publié: (2025)
Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies
par: Xiong, Tao, et autres
Publié: (2025)
par: Xiong, Tao, et autres
Publié: (2025)
When Babies Teach Babies: Can student knowledge sharing outperform Teacher-Guided Distillation on small datasets?
par: Iyer, Srikrishna
Publié: (2024)
par: Iyer, Srikrishna
Publié: (2024)
SMILE-Next: Teaching Large Language Models to Detect, Classify, and Reason about Laughter
par: Jung-Mok, Lee, et autres
Publié: (2026)
par: Jung-Mok, Lee, et autres
Publié: (2026)
Revisiting Anthropomorphic Reflection Markers in Large Language Model Reasoning
par: Yu, Yahan, et autres
Publié: (2026)
par: Yu, Yahan, et autres
Publié: (2026)
Reasoning Robustness of LLMs to Adversarial Typographical Errors
par: Gan, Esther, et autres
Publié: (2024)
par: Gan, Esther, et autres
Publié: (2024)
SciRIFF: A Resource to Enhance Language Model Instruction-Following over Scientific Literature
par: Wadden, David, et autres
Publié: (2024)
par: Wadden, David, et autres
Publié: (2024)
Backtracking Improves Generation Safety
par: Zhang, Yiming, et autres
Publié: (2024)
par: Zhang, Yiming, et autres
Publié: (2024)
Reinforcement Learning with Backtracking Feedback
par: Sel, Bilgehan, et autres
Publié: (2026)
par: Sel, Bilgehan, et autres
Publié: (2026)
Documents similaires
-
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
par: Kim, Joongwon, et autres
Publié: (2024) -
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
par: Kim, Joongwon, et autres
Publié: (2024) -
Data Engineering for Scaling Language Models to 128K Context
par: Fu, Yao, et autres
Publié: (2024) -
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models
par: Lyu, Xinxi, et autres
Publié: (2024) -
BTR: Binary Token Representations for Efficient Retrieval Augmented Language Models
par: Cao, Qingqing, et autres
Publié: (2023)