Saved in:
| Main Authors: | Twist, Lukas, Yannakoudakis, Helen, Zhang, Jie M. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.21127 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning
by: Twist, Lukas, et al.
Published: (2026)
by: Twist, Lukas, et al.
Published: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
by: Deshmukh, Pratik, et al.
Published: (2026)
by: Deshmukh, Pratik, et al.
Published: (2026)
Learning New Tasks from a Few Examples with Soft-Label Prototypes
by: Singh, Avyav Kumar, et al.
Published: (2022)
by: Singh, Avyav Kumar, et al.
Published: (2022)
A Functional Perspective on Knowledge Distillation in Neural Networks
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
by: Verhoeven, Ivo, et al.
Published: (2024)
by: Verhoeven, Ivo, et al.
Published: (2024)
A Study of LLMs' Preferences for Libraries and Programming Languages
by: Twist, Lukas, et al.
Published: (2025)
by: Twist, Lukas, et al.
Published: (2025)
A Function-Centric Perspective on Flat and Sharp Minima
by: Mason-Williams, Israel, et al.
Published: (2025)
by: Mason-Williams, Israel, et al.
Published: (2025)
RPO:Reinforcement Fine-Tuning with Partial Reasoning Optimization
by: Yi, Hongzhu, et al.
Published: (2026)
by: Yi, Hongzhu, et al.
Published: (2026)
RAGEN-2: Reasoning Collapse in Agentic RL
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
QuantLRM: Quantization of Large Reasoning Models via Fine-Tuning Signals
by: Zhang, Nan, et al.
Published: (2026)
by: Zhang, Nan, et al.
Published: (2026)
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
by: Zhao, Yanxiao, et al.
Published: (2025)
by: Zhao, Yanxiao, et al.
Published: (2025)
LIFT: Last-Mile Fine-Tuning for Table Explicitation
by: Khaitan, Divij, et al.
Published: (2026)
by: Khaitan, Divij, et al.
Published: (2026)
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
by: Springer, Max, et al.
Published: (2026)
by: Springer, Max, et al.
Published: (2026)
Reasoning Towards Fairness: Mitigating Bias in Language Models through Reasoning-Guided Fine-Tuning
by: Kabra, Sanchit, et al.
Published: (2025)
by: Kabra, Sanchit, et al.
Published: (2025)
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
by: Mu, Yongyu, et al.
Published: (2026)
by: Mu, Yongyu, et al.
Published: (2026)
Unsupervised Identification and Removal of Spurious Correlations During Fine-Tuning
by: Gilligan-Lee, Ciarán M., et al.
Published: (2026)
by: Gilligan-Lee, Ciarán M., et al.
Published: (2026)
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
by: Chen, Jinhao, et al.
Published: (2025)
by: Chen, Jinhao, et al.
Published: (2025)
ReasonIF: Large Reasoning Models Fail to Follow Instructions During Reasoning
by: Kwon, Yongchan, et al.
Published: (2025)
by: Kwon, Yongchan, et al.
Published: (2025)
Multimodal Fine-grained Reasoning for Post Quality Evaluation
by: Guo, Xiaoxu, et al.
Published: (2025)
by: Guo, Xiaoxu, et al.
Published: (2025)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
The Impact of Off-Policy Training Data on Probe Generalisation
by: Kirch, Nathalie, et al.
Published: (2025)
by: Kirch, Nathalie, et al.
Published: (2025)
Continual Calibration: Coverage Can Collapse Before Accuracy in Lifelong LLM Fine-Tuning
by: Shihab, Ibne Farabi, et al.
Published: (2026)
by: Shihab, Ibne Farabi, et al.
Published: (2026)
Learning to Trade Like an Expert: Cognitive Fine-Tuning for Stable Financial Reasoning in Language Models
by: Pan, Yuchen, et al.
Published: (2026)
by: Pan, Yuchen, et al.
Published: (2026)
One-Pass to Reason: Token Duplication and Block-Sparse Mask for Efficient Fine-Tuning on Multi-Turn Reasoning
by: Goru, Ritesh, et al.
Published: (2025)
by: Goru, Ritesh, et al.
Published: (2025)
Evaluating Mathematical Reasoning Across Large Language Models: A Fine-Grained Approach
by: Jahin, Afrar, et al.
Published: (2025)
by: Jahin, Afrar, et al.
Published: (2025)
Preventing Curriculum Collapse in Self-Evolving Reasoning Systems
by: Mishra, Vaibhav
Published: (2026)
by: Mishra, Vaibhav
Published: (2026)
Towards Robust Endogenous Reasoning: Unifying Drift Adaptation in Non-Stationary Tuning
by: Yang, Xiaoyu, et al.
Published: (2026)
by: Yang, Xiaoyu, et al.
Published: (2026)
Fine-Tuning Small Reasoning Models for Quantum Field Theory
by: Woodward, Nathaniel S., et al.
Published: (2026)
by: Woodward, Nathaniel S., et al.
Published: (2026)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
by: She, Jingyuan Selena, et al.
Published: (2023)
by: She, Jingyuan Selena, et al.
Published: (2023)
Directional Reasoning Trajectory Change (DRTC): Identifying Critical Trace Segments in Reasoning Models
by: Chang, Waldemar
Published: (2026)
by: Chang, Waldemar
Published: (2026)
Reasoning Model Unlearning: Forgetting Traces, Not Just Answers, While Preserving Reasoning Skills
by: Wang, Changsheng, et al.
Published: (2025)
by: Wang, Changsheng, et al.
Published: (2025)
PORT: Preference Optimization on Reasoning Traces
by: Lahlou, Salem, et al.
Published: (2024)
by: Lahlou, Salem, et al.
Published: (2024)
Stabilizing LLM Supervised Fine-Tuning via Explicit Distributional Control
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Fine-Tuning Without Forgetting via Loss-Adaptive Learning Rates
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
by: Prashant, Parjanya Prajakta, et al.
Published: (2026)
Criterion Collapse and Loss Distribution Control
by: Holland, Matthew J.
Published: (2024)
by: Holland, Matthew J.
Published: (2024)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Instruction Fine-Tuning: Does Prompt Loss Matter?
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
by: Huerta-Enochian, Mathew, et al.
Published: (2024)
When Safety Geometry Collapses: Fine-Tuning Vulnerabilities in Agentic Guard Models
by: Hossain, Ismail, et al.
Published: (2026)
by: Hossain, Ismail, et al.
Published: (2026)
Tracing Uncertainty in Language Model "Reasoning"
by: Grünefeld, Nils, et al.
Published: (2026)
by: Grünefeld, Nils, et al.
Published: (2026)
Similar Items
-
Not All Code Is Equal: A Data-Centric Study of Code Complexity and LLM Reasoning
by: Twist, Lukas, et al.
Published: (2026) -
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
by: Deshmukh, Pratik, et al.
Published: (2026) -
Learning New Tasks from a Few Examples with Soft-Label Prototypes
by: Singh, Avyav Kumar, et al.
Published: (2022) -
A Functional Perspective on Knowledge Distillation in Neural Networks
by: Mason-Williams, Israel, et al.
Published: (2025) -
A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
by: Verhoeven, Ivo, et al.
Published: (2024)