Introspective X Training: Feedback Conditioning Improves Scaling Across all LLM Training Stages
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Brandon, Lu, Ximing, Jung, Jaehun, Akter, Syeda Nahida, Kim, Hyunwoo, Qu, Yuxiao, Acuna, David, Prabhumoye, Shrimai, Choi, Yejin, Ammanabrolu, Prithviraj |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
by: Kim, Bosung, et al.
Published: (2026)
by: Kim, Bosung, et al.
Published: (2026)
RLP: Reinforcement as a Pretraining Objective
by: Hatamizadeh, Ali, et al.
Published: (2025)
by: Hatamizadeh, Ali, et al.
Published: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)
by: Lu, Ximing, et al.
Published: (2025)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
by: Jung, Jaehun, et al.
Published: (2025)
by: Jung, Jaehun, et al.
Published: (2025)
Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
by: Kim, Bosung, et al.
Published: (2025)
by: Kim, Bosung, et al.
Published: (2025)
iGRPO: Self-Feedback-Driven LLM Reasoning
by: Hatamizadeh, Ali, et al.
Published: (2026)
by: Hatamizadeh, Ali, et al.
Published: (2026)
Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
by: Lu, Ximing, et al.
Published: (2026)
by: Lu, Ximing, et al.
Published: (2026)
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Nemotron-CrossThink: Scaling Self-Learning beyond Math Reasoning
by: Akter, Syeda Nahida, et al.
Published: (2025)
by: Akter, Syeda Nahida, et al.
Published: (2025)
Socratic-MCTS: Test-Time Visual Reasoning by Asking the Right Questions
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
How Reasoning Evolves from Post-Training Data: An Empirical Study Using Chess
by: Dionisopoulos, Lucas, et al.
Published: (2026)
by: Dionisopoulos, Lucas, et al.
Published: (2026)
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
by: Wang, Ruiyi, et al.
Published: (2025)
by: Wang, Ruiyi, et al.
Published: (2025)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
by: Shen, Yiran, et al.
Published: (2025)
by: Shen, Yiran, et al.
Published: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Critique-out-Loud Reward Models
by: Ankner, Zachary, et al.
Published: (2024)
by: Ankner, Zachary, et al.
Published: (2024)
JAMDEC: Unsupervised Authorship Obfuscation using Constrained Decoding over Small Language Models
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight
by: Cui, Christopher Z., et al.
Published: (2026)
by: Cui, Christopher Z., et al.
Published: (2026)
Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
by: Tasnim, Nazia, et al.
Published: (2026)
by: Tasnim, Nazia, et al.
Published: (2026)
Synthetic Mixed Training: Scaling Parametric Knowledge Acquisition Beyond RAG
by: Han, Seungju, et al.
Published: (2026)
by: Han, Seungju, et al.
Published: (2026)
Information-Theoretic Distillation for Reference-less Summarization
by: Jung, Jaehun, et al.
Published: (2024)
by: Jung, Jaehun, et al.
Published: (2024)
Preference-Based Learning in Audio Applications: A Systematic Analysis
by: Broukhim, Aaron, et al.
Published: (2025)
by: Broukhim, Aaron, et al.
Published: (2025)
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
by: Jung, Jaehun, et al.
Published: (2023)
by: Jung, Jaehun, et al.
Published: (2023)
Recursive Introspection: Teaching Language Model Agents How to Self-Improve
by: Qu, Yuxiao, et al.
Published: (2024)
by: Qu, Yuxiao, et al.
Published: (2024)
TALES: Text Adventure Learning Environment Suite
by: Cui, Christopher Zhang, et al.
Published: (2025)
by: Cui, Christopher Zhang, et al.
Published: (2025)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
by: Seo, Wooseok, et al.
Published: (2025)
by: Seo, Wooseok, et al.
Published: (2025)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
by: Lee, Jaeyoung, et al.
Published: (2024)
by: Lee, Jaeyoung, et al.
Published: (2024)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Introspection Adapters: Training LLMs to Report Their Learned Behaviors
by: Shenoy, Keshav, et al.
Published: (2026)
by: Shenoy, Keshav, et al.
Published: (2026)
Nemotron-CC-Math: A 133 Billion-Token-Scale High Quality Math Pretraining Dataset
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
by: Mahabadi, Rabeeh Karimi, et al.
Published: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
by: Hallinan, Skyler, et al.
Published: (2025)
by: Hallinan, Skyler, et al.
Published: (2025)
CPS-TaskForge: Generating Collaborative Problem Solving Environments for Diverse Communication Tasks
by: Haduong, Nikita, et al.
Published: (2024)
by: Haduong, Nikita, et al.
Published: (2024)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
by: Ravichander, Abhilasha, et al.
Published: (2025)
by: Ravichander, Abhilasha, et al.
Published: (2025)
Training neural control variates using correlated configurations
by: Oh, Hyunwoo
Published: (2025)
by: Oh, Hyunwoo
Published: (2025)
Maximize Your Data's Potential: Enhancing LLM Accuracy with Two-Phase Pretraining
by: Feng, Steven, et al.
Published: (2024)
by: Feng, Steven, et al.
Published: (2024)
A Study on Scaling Up Multilingual News Framing Analysis
by: Akter, Syeda Sabrina, et al.
Published: (2024)
by: Akter, Syeda Sabrina, et al.
Published: (2024)
Learning Context-Conditioned Predicate Semantics via Prototype Feedback
by: Jung, NamGyu, et al.
Published: (2026)
by: Jung, NamGyu, et al.
Published: (2026)
Similar Items
-
Front-Loading Reasoning: The Synergy between Pretraining and Post-Training Data
by: Akter, Syeda Nahida, et al.
Published: (2025) -
How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning
by: Kim, Bosung, et al.
Published: (2026) -
RLP: Reinforcement as a Pretraining Objective
by: Hatamizadeh, Ali, et al.
Published: (2025) -
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025) -
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning
by: Lu, Ximing, et al.
Published: (2025)