Predicting LLM Reasoning Performance with Small Proxy Model
Fuente:
arXiv
Saved in:
| Main Authors: | Koh, Woosung, Suk, Juyoung, Han, Sungjun, Yun, Se-Young, Shin, Jamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Trillion 7B Technical Report
by: Han, Sungjun, et al.
Published: (2025)
by: Han, Sungjun, et al.
Published: (2025)
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Extreme Solar Flare Prediction Using Residual Networks with HMI Magnetograms and Intensitygrams
by: Yun, Juyoung, et al.
Published: (2024)
by: Yun, Juyoung, et al.
Published: (2024)
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Analysis and Predictive Modeling of Solar Coronal Holes Using Computer Vision and ARIMA-LSTM Networks
by: Yun, Juyoung, et al.
Published: (2024)
by: Yun, Juyoung, et al.
Published: (2024)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Mitigating Gradient Overlap in Deep Residual Networks with Gradient Normalization for Improved Non-Convex Optimization
by: Yun, Juyoung
Published: (2024)
by: Yun, Juyoung
Published: (2024)
AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
What is the Alignment Objective of GRPO?
by: Vojnovic, Milan, et al.
Published: (2025)
by: Vojnovic, Milan, et al.
Published: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Temporal Alignment Guidance: On-Manifold Sampling in Diffusion Models
by: Park, Youngrok, et al.
Published: (2025)
by: Park, Youngrok, et al.
Published: (2025)
ProxyKV: Cross-Model Proxy Pruning for Efficient Long-Context LLM Inference
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
Encoding Temporal Statistical-space Priors via Augmented Representation
by: Choi, Insu, et al.
Published: (2024)
by: Choi, Insu, et al.
Published: (2024)
Curriculum Learning and Imitation Learning for Model-free Control on Financial Time-series
by: Koh, Woosung, et al.
Published: (2023)
by: Koh, Woosung, et al.
Published: (2023)
Self-Refinement of Language Models from External Proxy Metrics Feedback
by: Ramji, Keshav, et al.
Published: (2024)
by: Ramji, Keshav, et al.
Published: (2024)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
The Hidden Power of Pure 16-bit Floating-Point Neural Networks
by: Yun, Juyoung, et al.
Published: (2023)
by: Yun, Juyoung, et al.
Published: (2023)
Stochastic Gradient Sampling for Enhancing Neural Networks Training
by: Yun, Juyoung
Published: (2023)
by: Yun, Juyoung
Published: (2023)
RetroReasoner: A Reasoning LLM for Strategic Retrosynthesis Prediction
by: Ko, Hanbum, et al.
Published: (2026)
by: Ko, Hanbum, et al.
Published: (2026)
Contextual Linear Bandits under Noisy Features: Towards Bayesian Oracles
by: Kim, Jung-hun, et al.
Published: (2017)
by: Kim, Jung-hun, et al.
Published: (2017)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
by: Wang, Jiachen T., et al.
Published: (2025)
by: Wang, Jiachen T., et al.
Published: (2025)
Conditional Synthesis of 3D Molecules with Time Correction Sampler
by: Jung, Hojung, et al.
Published: (2024)
by: Jung, Hojung, et al.
Published: (2024)
A Hierarchical Language Model with Predictable Scaling Laws and Provable Benefits of Reasoning
by: Gaitonde, Jason, et al.
Published: (2026)
by: Gaitonde, Jason, et al.
Published: (2026)
DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
by: Ko, Jongwoo, et al.
Published: (2025)
by: Ko, Jongwoo, et al.
Published: (2025)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
by: Shin, Kwan Soo
Published: (2026)
by: Shin, Kwan Soo
Published: (2026)
Dynamics-Predictive Sampling for Active RL Finetuning of Large Reasoning Models
by: Mao, Yixiu, et al.
Published: (2026)
by: Mao, Yixiu, et al.
Published: (2026)
Automated Filtering of Human Feedback Data for Aligning Text-to-Image Diffusion Models
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning
by: Kim, Hyeonjin, et al.
Published: (2026)
by: Kim, Hyeonjin, et al.
Published: (2026)
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness
by: Yang, Yongjin, et al.
Published: (2025)
by: Yang, Yongjin, et al.
Published: (2025)
Automated Skill Discovery for Language Agents through Exploration and Iterative Feedback
by: Yang, Yongjin, et al.
Published: (2025)
by: Yang, Yongjin, et al.
Published: (2025)
Towards Mitigating Excessive Forgetting in LLM Unlearning via Entanglement-Guidance with Proxy Constraint
by: Liu, Zhihao, et al.
Published: (2025)
by: Liu, Zhihao, et al.
Published: (2025)
Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
by: Qu, Yun, et al.
Published: (2025)
by: Qu, Yun, et al.
Published: (2025)
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search
by: Han, Dongge, et al.
Published: (2025)
by: Han, Dongge, et al.
Published: (2025)
Proxy-RLHF: Decoupling Generation and Alignment in Large Language Model with Proxy
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
FairDICE: Fairness-Driven Offline Multi-Objective Reinforcement Learning
by: Kim, Woosung, et al.
Published: (2025)
by: Kim, Woosung, et al.
Published: (2025)
RAST: Reasoning Activation in LLMs via Small-model Transfer
by: Ouyang, Siru, et al.
Published: (2025)
by: Ouyang, Siru, et al.
Published: (2025)
Lightweight Dataset Pruning without Full Training via Example Difficulty and Prediction Uncertainty
by: Cho, Yeseul, et al.
Published: (2025)
by: Cho, Yeseul, et al.
Published: (2025)
CaliciBoost: Performance-Driven Evaluation of Molecular Representations for Caco-2 Permeability Prediction
by: Van Le, Huong, et al.
Published: (2025)
by: Van Le, Huong, et al.
Published: (2025)
Similar Items
-
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026) -
Trillion 7B Technical Report
by: Han, Sungjun, et al.
Published: (2025) -
mSFT: Addressing Dataset Mixtures Overfitting Heterogeneously in Multi-task SFT
by: Koh, Woosung, et al.
Published: (2026) -
Extreme Solar Flare Prediction Using Residual Networks with HMI Magnetograms and Intensitygrams
by: Yun, Juyoung, et al.
Published: (2024) -
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
by: Koh, Woosung, et al.
Published: (2024)