Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Minwu, Shrestha, Safal, Shrestha, Anubhav, Ross, Keith |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
por: Shrestha, Safal, et al.
Publicado: (2025)
por: Shrestha, Safal, et al.
Publicado: (2025)
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
por: Shrestha, Safal, et al.
Publicado: (2025)
por: Shrestha, Safal, et al.
Publicado: (2025)
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
por: Shrestha, Safal, et al.
Publicado: (2026)
por: Shrestha, Safal, et al.
Publicado: (2026)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
por: Kim, Minwu, et al.
Publicado: (2025)
por: Kim, Minwu, et al.
Publicado: (2025)
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
por: Nepal, Aadim, et al.
Publicado: (2025)
por: Nepal, Aadim, et al.
Publicado: (2025)
Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes
por: Setlur, Amrith, et al.
Publicado: (2026)
por: Setlur, Amrith, et al.
Publicado: (2026)
Token-Driven GammaTune: Adaptive Calibration for Enhanced Speculative Decoding
por: Gautam, Aayush, et al.
Publicado: (2025)
por: Gautam, Aayush, et al.
Publicado: (2025)
Deja vu: Contrastive Historical Modeling with Prefix-tuning for Temporal Knowledge Graph Reasoning
por: Peng, Miao, et al.
Publicado: (2024)
por: Peng, Miao, et al.
Publicado: (2024)
Large Language Model Reasoning Failures
por: Song, Peiyang, et al.
Publicado: (2026)
por: Song, Peiyang, et al.
Publicado: (2026)
Structured Extraction of Real World Medical Knowledge using LLMs for Summarization and Search
por: Kim, Edward, et al.
Publicado: (2024)
por: Kim, Edward, et al.
Publicado: (2024)
Query-Conditioned Test-Time Self-Training for Large Language Models
por: Song, Chaehee, et al.
Publicado: (2026)
por: Song, Chaehee, et al.
Publicado: (2026)
Towards Infinite-Long Prefix in Transformer
por: Liang, Yingyu, et al.
Publicado: (2024)
por: Liang, Yingyu, et al.
Publicado: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
por: Munkhbat, Tergel, et al.
Publicado: (2025)
por: Munkhbat, Tergel, et al.
Publicado: (2025)
RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems
por: Qu, Yuxiao, et al.
Publicado: (2025)
por: Qu, Yuxiao, et al.
Publicado: (2025)
Distribution Matching via Generalized Consistency Models
por: Shrestha, Sagar, et al.
Publicado: (2025)
por: Shrestha, Sagar, et al.
Publicado: (2025)
Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting
por: Shrestha, Gyanendra, et al.
Publicado: (2025)
por: Shrestha, Gyanendra, et al.
Publicado: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
por: Huang, Zeyu, et al.
Publicado: (2025)
por: Huang, Zeyu, et al.
Publicado: (2025)
Efficient Multi-Hop Question Answering over Knowledge Graphs via LLM Planning and Embedding-Guided Search
por: Shrestha, Manil, et al.
Publicado: (2025)
por: Shrestha, Manil, et al.
Publicado: (2025)
L-TUNING: Synchronized Label Tuning for Prompt and Prefix in LLMs
por: Kowsher, Md., et al.
Publicado: (2023)
por: Kowsher, Md., et al.
Publicado: (2023)
The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity
por: Shojaee, Parshin, et al.
Publicado: (2025)
por: Shojaee, Parshin, et al.
Publicado: (2025)
Semantic Refinement with LLMs for Graph Representations
por: Thapaliya, Safal, et al.
Publicado: (2025)
por: Thapaliya, Safal, et al.
Publicado: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
por: Xi, Zhiheng, et al.
Publicado: (2024)
por: Xi, Zhiheng, et al.
Publicado: (2024)
HAPO: Training Language Models to Reason Concisely via History-Aware Policy Optimization
por: Huang, Chengyu, et al.
Publicado: (2025)
por: Huang, Chengyu, et al.
Publicado: (2025)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
por: Koh, Woosung, et al.
Publicado: (2025)
por: Koh, Woosung, et al.
Publicado: (2025)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
por: Ballon, Marthe, et al.
Publicado: (2026)
por: Ballon, Marthe, et al.
Publicado: (2026)
On the Optimal Reasoning Length for RL-Trained Language Models
por: Nohara, Daisuke, et al.
Publicado: (2026)
por: Nohara, Daisuke, et al.
Publicado: (2026)
POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration
por: Qu, Yuxiao, et al.
Publicado: (2026)
por: Qu, Yuxiao, et al.
Publicado: (2026)
Nwāchā Munā: A Devanagari Speech Corpus and Proximal Transfer Benchmark for Nepal Bhasha ASR
por: Sharma, Rishikesh Kumar, et al.
Publicado: (2026)
por: Sharma, Rishikesh Kumar, et al.
Publicado: (2026)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
por: Kang, Deokhyung, et al.
Publicado: (2025)
por: Kang, Deokhyung, et al.
Publicado: (2025)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
por: Damani, Mehul, et al.
Publicado: (2025)
por: Damani, Mehul, et al.
Publicado: (2025)
Can A Gamer Train A Mathematical Reasoning Model?
por: Shin, Andrew
Publicado: (2025)
por: Shin, Andrew
Publicado: (2025)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
por: Kim, Donghoon, et al.
Publicado: (2024)
por: Kim, Donghoon, et al.
Publicado: (2024)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
por: Bai, Yuyang, et al.
Publicado: (2026)
por: Bai, Yuyang, et al.
Publicado: (2026)
Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models
por: Dang, Trung Cuong, et al.
Publicado: (2025)
por: Dang, Trung Cuong, et al.
Publicado: (2025)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
por: Liu, Mingjie, et al.
Publicado: (2025)
por: Liu, Mingjie, et al.
Publicado: (2025)
IDAT: A Multi-Modal Dataset and Toolkit for Building and Evaluating Interactive Task-Solving Agents
por: Mohanty, Shrestha, et al.
Publicado: (2024)
por: Mohanty, Shrestha, et al.
Publicado: (2024)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
por: Sahoo, Subramanyam, et al.
Publicado: (2026)
VLSM-Adapter: Finetuning Vision-Language Segmentation Efficiently with Lightweight Blocks
por: Dhakal, Manish, et al.
Publicado: (2024)
por: Dhakal, Manish, et al.
Publicado: (2024)
Training Large Language Models To Reason In Parallel With Global Forking Tokens
por: Jia, Sheng, et al.
Publicado: (2025)
por: Jia, Sheng, et al.
Publicado: (2025)
Ejemplares similares
-
Mathematical Reasoning in Large Language Models: Assessing Logical and Arithmetic Errors across Wide Numerical Ranges
por: Shrestha, Safal, et al.
Publicado: (2025) -
Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings
por: Shrestha, Safal, et al.
Publicado: (2025) -
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
por: Shrestha, Safal, et al.
Publicado: (2026) -
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
por: Kim, Minwu, et al.
Publicado: (2025) -
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training
por: Nepal, Aadim, et al.
Publicado: (2025)