A Primer in Post-Training Reasoning Data: What We Know About How It Works
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yaoming, Zhao, Guangxiang, Shi, Qilong, Sun, Lin, Zhang, Xiangzheng, Yang, Tong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
di: Zhao, Guangxiang, et al.
Pubblicazione: (2026)
di: Zhao, Guangxiang, et al.
Pubblicazione: (2026)
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
di: Zhu, Yongfu, et al.
Pubblicazione: (2025)
di: Zhu, Yongfu, et al.
Pubblicazione: (2025)
Beyond Parameter Arithmetic: Sparse Complementary Fusion for Distribution-Aware Model Merging
di: Lin, Weihong, et al.
Pubblicazione: (2026)
di: Lin, Weihong, et al.
Pubblicazione: (2026)
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
di: Si, Jianfeng, et al.
Pubblicazione: (2025)
di: Si, Jianfeng, et al.
Pubblicazione: (2025)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
di: Zhao, Guangxiang, et al.
Pubblicazione: (2025)
di: Zhao, Guangxiang, et al.
Pubblicazione: (2025)
Thinking About Thinking: Evaluating Reasoning in Post-Trained Language Models
di: Singla, Pratham, et al.
Pubblicazione: (2025)
di: Singla, Pratham, et al.
Pubblicazione: (2025)
What We Talk About When We Talk About LMs: Implicit Paradigm Shifts and the Ship of Language Models
di: Zhu, Shengqi, et al.
Pubblicazione: (2024)
di: Zhu, Shengqi, et al.
Pubblicazione: (2024)
NanoKnow: How to Know What Your Language Model Knows
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
di: Gu, Lingwei, et al.
Pubblicazione: (2026)
LLMs for Relational Reasoning: How Far are We?
di: Li, Zhiming, et al.
Pubblicazione: (2024)
di: Li, Zhiming, et al.
Pubblicazione: (2024)
TinyR1-32B-Preview: Boosting Accuracy with Branch-Merge Distillation
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Should We be Pedantic About Reasoning Errors in Machine Translation?
di: Bao, Calvin, et al.
Pubblicazione: (2026)
di: Bao, Calvin, et al.
Pubblicazione: (2026)
Can AI Assistants Know What They Don't Know?
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
di: Cheng, Qinyuan, et al.
Pubblicazione: (2024)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
di: Lee, Joosung, et al.
Pubblicazione: (2026)
di: Lee, Joosung, et al.
Pubblicazione: (2026)
Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
di: Yao, Yilun, et al.
Pubblicazione: (2026)
di: Yao, Yilun, et al.
Pubblicazione: (2026)
KnowRL: Teaching Language Models to Know What They Know
di: Kale, Sahil, et al.
Pubblicazione: (2025)
di: Kale, Sahil, et al.
Pubblicazione: (2025)
How Far Are We from Optimal Reasoning Efficiency?
di: Gao, Jiaxuan, et al.
Pubblicazione: (2025)
di: Gao, Jiaxuan, et al.
Pubblicazione: (2025)
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method
di: Zhao, Yukun, et al.
Pubblicazione: (2023)
di: Zhao, Yukun, et al.
Pubblicazione: (2023)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
di: Jiang, Lei, et al.
Pubblicazione: (2026)
di: Jiang, Lei, et al.
Pubblicazione: (2026)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
di: Mei, Zhiting, et al.
Pubblicazione: (2025)
Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
di: Lin, Wenye, et al.
Pubblicazione: (2026)
di: Lin, Wenye, et al.
Pubblicazione: (2026)
BertaQA: How Much Do Language Models Know About Local Culture?
di: Etxaniz, Julen, et al.
Pubblicazione: (2024)
di: Etxaniz, Julen, et al.
Pubblicazione: (2024)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
di: Ceron, Tanise, et al.
Pubblicazione: (2025)
di: Ceron, Tanise, et al.
Pubblicazione: (2025)
Detecting RLVR Training Data via Structural Convergence of Reasoning
di: Zhang, Hongbo, et al.
Pubblicazione: (2026)
di: Zhang, Hongbo, et al.
Pubblicazione: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
di: Lv, Keyu, et al.
Pubblicazione: (2026)
di: Lv, Keyu, et al.
Pubblicazione: (2026)
Do Retrieval Augmented Language Models Know When They Don't Know?
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
di: Zhou, Youchao, et al.
Pubblicazione: (2025)
Dense X Retrieval: What Retrieval Granularity Should We Use?
di: Chen, Tong, et al.
Pubblicazione: (2023)
di: Chen, Tong, et al.
Pubblicazione: (2023)
KnowCoder-A1: Incentivizing Agentic Reasoning Capability with Outcome Supervision for KBQA
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
di: Chen, Zhuo, et al.
Pubblicazione: (2025)
Reasoning Models Will Sometimes Lie About Their Reasoning
di: Walden, William, et al.
Pubblicazione: (2026)
di: Walden, William, et al.
Pubblicazione: (2026)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
di: Damani, Mehul, et al.
Pubblicazione: (2025)
di: Damani, Mehul, et al.
Pubblicazione: (2025)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
di: Wang, Chenyang, et al.
Pubblicazione: (2025)
SnapKV: LLM Knows What You are Looking for Before Generation
di: Li, Yuhong, et al.
Pubblicazione: (2024)
di: Li, Yuhong, et al.
Pubblicazione: (2024)
Do Large Language Models Know What They Are Capable Of?
di: Barkan, Casey O., et al.
Pubblicazione: (2025)
di: Barkan, Casey O., et al.
Pubblicazione: (2025)
From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
di: Wang, Shaojie, et al.
Pubblicazione: (2026)
di: Wang, Shaojie, et al.
Pubblicazione: (2026)
Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoning
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
di: Zhao, Qiannian, et al.
Pubblicazione: (2026)
An Efficient and Precise Training Data Construction Framework for Process-supervised Reward Model in Mathematical Reasoning
di: Sun, Wei, et al.
Pubblicazione: (2025)
di: Sun, Wei, et al.
Pubblicazione: (2025)
Language Agents Mirror Human Causal Reasoning Biases. How Can We Help Them Think Like Scientists?
di: GX-Chen, Anthony, et al.
Pubblicazione: (2025)
di: GX-Chen, Anthony, et al.
Pubblicazione: (2025)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
di: Liu, Yi, et al.
Pubblicazione: (2025)
di: Liu, Yi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Thinking with Reasoning Skills: Fewer Tokens, More Accuracy
di: Zhao, Guangxiang, et al.
Pubblicazione: (2026) -
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM
di: Zhu, Yongfu, et al.
Pubblicazione: (2025) -
Beyond Parameter Arithmetic: Sparse Complementary Fusion for Distribution-Aware Model Merging
di: Lin, Weihong, et al.
Pubblicazione: (2026) -
Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training
di: Si, Jianfeng, et al.
Pubblicazione: (2025) -
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)