Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Yihao, Zhang, Allan, Huang, Jianhao, Sahai, Amit, Mirzasoleiman, Baharan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
LoRA is All You Need for Safety Alignment of Reasoning LLMs
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Investigating the Impact of Model Width and Density on Generalization in Presence of Label Noise
von: Xue, Yihao, et al.
Veröffentlicht: (2022)
von: Xue, Yihao, et al.
Veröffentlicht: (2022)
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Dataset Distillation via Knowledge Distillation: Towards Efficient Self-Supervised Pre-Training of Deep Networks
von: Joshi, Siddharth, et al.
Veröffentlicht: (2024)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2024)
Data-Efficient Contrastive Self-supervised Learning: Most Beneficial Examples for Supervised Learning Contribute the Least
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
Graph Contrastive Learning under Heterophily via Graph Filters
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
von: Javanmard, Adel, et al.
Veröffentlicht: (2026)
Few-shot Adaptation to Distribution Shifts By Mixing Source and Target Embeddings
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
von: Xue, Yihao, et al.
Veröffentlicht: (2023)
Challenges and Opportunities in Improving Worst-Group Generalization in Presence of Spurious Features
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2023)
Beyond Semantic Entropy: Boosting LLM Uncertainty Quantification with Pairwise Semantic Similarity
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
von: Xue, Yihao, et al.
Veröffentlicht: (2025)
Investigating the Benefits of Projection Head for Representation Learning
von: Xue, Yihao, et al.
Veröffentlicht: (2024)
von: Xue, Yihao, et al.
Veröffentlicht: (2024)
How Transformers Learn to Plan via Multi-Token Prediction
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
von: Huang, Jianhao, et al.
Veröffentlicht: (2026)
Changing the Training Data Distribution to Reduce Simplicity Bias Improves In-distribution Generalization
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Identifying Spurious Biases Early in Training through the Lens of Simplicity Bias
von: Yang, Yu, et al.
Veröffentlicht: (2023)
von: Yang, Yu, et al.
Veröffentlicht: (2023)
Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor Attacks
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
von: Yang, Wenhan, et al.
Veröffentlicht: (2023)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models
von: Yang, Yu, et al.
Veröffentlicht: (2024)
von: Yang, Yu, et al.
Veröffentlicht: (2024)
Mini-batch Coresets for Memory-efficient Language Model Training on Data Mixtures
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
von: Nguyen, Dang, et al.
Veröffentlicht: (2024)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
von: Joshi, Siddharth, et al.
Veröffentlicht: (2024)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2024)
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Synthetic Text Generation for Training Large Language Models via Gradient Matching
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Data Distribution as a Lever for Guiding Optimizers Toward Superior Generalization in LLMs
von: Gangavarapu, Tushaar, et al.
Veröffentlicht: (2026)
von: Gangavarapu, Tushaar, et al.
Veröffentlicht: (2026)
Efficient Reasoning at Fixed Test-Time Cost via Length-Aware Attention Priors and Gain-Aware Training
von: Atri, Rian
Veröffentlicht: (2026)
von: Atri, Rian
Veröffentlicht: (2026)
Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
von: Naharas, Nilay, et al.
Veröffentlicht: (2025)
Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
von: Ahmadpour, Mohammadjavad, et al.
Veröffentlicht: (2025)
Intent Laundering: AI Safety Datasets Are Not What They Seem
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
von: Golchin, Shahriar, et al.
Veröffentlicht: (2026)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
von: Joshi, Siddharth, et al.
Veröffentlicht: (2025)
Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap
von: Yang, Wenhan, et al.
Veröffentlicht: (2025)
von: Yang, Wenhan, et al.
Veröffentlicht: (2025)
NeWRF: A Deep Learning Framework for Wireless Radiation Field Reconstruction and Channel Prediction
von: Lu, Haofan, et al.
Veröffentlicht: (2024)
von: Lu, Haofan, et al.
Veröffentlicht: (2024)
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models
von: Phan, Hoang, et al.
Veröffentlicht: (2025)
von: Phan, Hoang, et al.
Veröffentlicht: (2025)
What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
von: Feng, Yunzhen, et al.
Veröffentlicht: (2025)
von: Feng, Yunzhen, et al.
Veröffentlicht: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
Test Time Training for Supervised Causal Learning
von: Deng, Zizhen, et al.
Veröffentlicht: (2026)
von: Deng, Zizhen, et al.
Veröffentlicht: (2026)
Multi-Turn Jailbreaks Are Simpler Than They Seem
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoxue, et al.
Veröffentlicht: (2025)
Predictive Auditing of Hidden Tokens in LLM APIs via Reasoning Length Estimation
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
von: Wang, Ziyao, et al.
Veröffentlicht: (2025)
Sufficient and Necessary Explanations (and What Lies in Between)
von: Bharti, Beepul, et al.
Veröffentlicht: (2024)
von: Bharti, Beepul, et al.
Veröffentlicht: (2024)
Beyond Correctness: Harmonizing Process and Outcome Rewards through RL Training
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
von: Ye, Chenlu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from $k$-Parity
von: Huang, Jianhao, et al.
Veröffentlicht: (2026) -
LoRA is All You Need for Safety Alignment of Reasoning LLMs
von: Xue, Yihao, et al.
Veröffentlicht: (2025) -
Investigating the Impact of Model Width and Density on Generalization in Presence of Label Noise
von: Xue, Yihao, et al.
Veröffentlicht: (2022) -
Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
von: Xue, Yihao, et al.
Veröffentlicht: (2025) -
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)