Reasoning Steps as Curriculum: Using Depth of Thought as a Difficulty Signal for Tuning LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Jung, Jeesu, Jung, Sangkeun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024)
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2025)
di: Heo, Inbum, et al.
Pubblicazione: (2025)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2026)
di: Heo, Inbum, et al.
Pubblicazione: (2026)
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
di: Hwang, Taewook, et al.
Pubblicazione: (2024)
Phasor Memory Networks: Stable Backpropagation Through Time for Scalable Explicit Memory
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
di: Goo, Sungwoo, et al.
Pubblicazione: (2026)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
di: Fernandez, Nigel, et al.
Pubblicazione: (2025)
Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
di: Kang, Hyeonseok, et al.
Pubblicazione: (2024)
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
di: Tzannetos, Georgios, et al.
Pubblicazione: (2025)
di: Tzannetos, Georgios, et al.
Pubblicazione: (2025)
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
di: Mahrooghi, Ilia, et al.
Pubblicazione: (2026)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
MathMixup: Boosting LLM Mathematical Reasoning with Difficulty-Controllable Data Synthesis and Curriculum Learning
di: Li, Xuchen, et al.
Pubblicazione: (2026)
di: Li, Xuchen, et al.
Pubblicazione: (2026)
Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
di: Di, Xinhan, et al.
Pubblicazione: (2025)
di: Di, Xinhan, et al.
Pubblicazione: (2025)
The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains
di: Kang, Jung Min
Pubblicazione: (2026)
di: Kang, Jung Min
Pubblicazione: (2026)
Thought Anchors: Which LLM Reasoning Steps Matter?
di: Bogdan, Paul C., et al.
Pubblicazione: (2025)
di: Bogdan, Paul C., et al.
Pubblicazione: (2025)
Demystifying Long Chain-of-Thought Reasoning in LLMs
di: Yeo, Edward, et al.
Pubblicazione: (2025)
di: Yeo, Edward, et al.
Pubblicazione: (2025)
Boosting Deductive Reasoning with Step Signals In RLHF
di: Li, Jialian, et al.
Pubblicazione: (2024)
di: Li, Jialian, et al.
Pubblicazione: (2024)
Does the Definition of Difficulty Matter? Scoring Functions and their Role for Curriculum Learning
di: Rampp, Simon, et al.
Pubblicazione: (2024)
di: Rampp, Simon, et al.
Pubblicazione: (2024)
When Thoughts Meet Facts: Reusable Reasoning for Long-Context LMs
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
di: Jeong, Soyeong, et al.
Pubblicazione: (2025)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic
di: Mao, Zhenjiang, et al.
Pubblicazione: (2025)
di: Mao, Zhenjiang, et al.
Pubblicazione: (2025)
DAG-Math: Graph-of-Thought Guided Mathematical Reasoning in LLMs
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
di: Zhang, Yuanhe, et al.
Pubblicazione: (2025)
ARIES: Autonomous Reasoning with LLMs on Interactive Thought Graph Environments
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
di: Gimenes, Pedro, et al.
Pubblicazione: (2025)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
di: Zhang, Xuan, et al.
Pubblicazione: (2024)
Evaluating LLMs' Reasoning Over Ordered Procedural Steps
di: Anika, Adrita, et al.
Pubblicazione: (2025)
di: Anika, Adrita, et al.
Pubblicazione: (2025)
TopicTag: Automatic Annotation of NMF Topic Models Using Chain of Thought and Prompt Tuning with LLMs
di: Wanna, Selma, et al.
Pubblicazione: (2024)
di: Wanna, Selma, et al.
Pubblicazione: (2024)
Denoising Task Difficulty-based Curriculum for Training Diffusion Models
di: Kim, Jin-Young, et al.
Pubblicazione: (2024)
di: Kim, Jin-Young, et al.
Pubblicazione: (2024)
A Survey on Hypergraph Neural Networks: An In-Depth and Step-By-Step Guide
di: Kim, Sunwoo, et al.
Pubblicazione: (2024)
di: Kim, Sunwoo, et al.
Pubblicazione: (2024)
Unifying Block-wise PTQ and Distillation-based QAT for Progressive Quantization toward 2-bit Instruction-Tuned LLMs
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
di: Lee, Jung Hyun, et al.
Pubblicazione: (2025)
Understanding the Difficulty of Low-Precision Post-Training Quantization for LLMs
di: Xu, Zifei, et al.
Pubblicazione: (2024)
di: Xu, Zifei, et al.
Pubblicazione: (2024)
Long-Context Reasoning Through Proxy-Based Chain-of-Thought Tuning
di: Li, Miao, et al.
Pubblicazione: (2026)
di: Li, Miao, et al.
Pubblicazione: (2026)
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
di: Lai, Xin, et al.
Pubblicazione: (2024)
di: Lai, Xin, et al.
Pubblicazione: (2024)
Hidden Error Awareness in Chain-of-Thought Reasoning: The Signal Is Diagnostic, Not Causal
di: Yuan, Aojie, et al.
Pubblicazione: (2026)
di: Yuan, Aojie, et al.
Pubblicazione: (2026)
Shorter Thoughts, Same Answers: Difficulty-Scaled Segment-Wise RL for CoT Compression
di: Tian, Ye, et al.
Pubblicazione: (2026)
di: Tian, Ye, et al.
Pubblicazione: (2026)
Data Difficulty and the Generalization--Extrapolation Tradeoff in LLM Fine-Tuning
di: Liu, Siyuan, et al.
Pubblicazione: (2026)
di: Liu, Siyuan, et al.
Pubblicazione: (2026)
CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs
di: Zeng, Yongcheng, et al.
Pubblicazione: (2025)
di: Zeng, Yongcheng, et al.
Pubblicazione: (2025)
ABench-Physics: Benchmarking Physical Reasoning in LLMs via High-Difficulty and Dynamic Physics Problems
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
di: Zhang, Yiming, et al.
Pubblicazione: (2025)
Understanding Sharpness Dynamics in NN Training with a Minimalist Example: The Effects of Dataset Difficulty, Depth, Stochasticity, and More
di: Yoo, Geonhui, et al.
Pubblicazione: (2025)
di: Yoo, Geonhui, et al.
Pubblicazione: (2025)
Masked Thought: Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language Models
di: Chen, Changyu, et al.
Pubblicazione: (2024)
di: Chen, Changyu, et al.
Pubblicazione: (2024)
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
di: Zhao, Chengshuai, et al.
Pubblicazione: (2025)
di: Zhao, Chengshuai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Exploring Domain Robust Lightweight Reward Models based on Router Mechanism
di: Namgoong, Hyuk, et al.
Pubblicazione: (2024) -
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025) -
LED Benchmark: Diagnosing Structural Layout Errors for Document Layout Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2025) -
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
di: Heo, Inbum, et al.
Pubblicazione: (2026) -
Employing Layerwised Unsupervised Learning to Lessen Data and Loss Requirements in Forward-Forward Algorithms
di: Hwang, Taewook, et al.
Pubblicazione: (2024)