AdaSTaR: Adaptive Data Sampling for Training Self-Taught Reasoners
Fuente:
arXiv
Saved in:
| Main Authors: | Koh, Woosung, Oh, Wonbeen, Jang, Jaein, Lee, MinHyung, Kim, Hyeongjin, Kim, Ah Yeon, Kim, Joonkee, Lee, Junghyun, Kim, Taehyeon, Yun, Se-Young |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
by: Kim, Taehyeon, et al.
Published: (2023)
by: Kim, Taehyeon, et al.
Published: (2023)
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024)
by: Koh, Woosung, et al.
Published: (2024)
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
by: Kim, Taehyeon, et al.
Published: (2024)
by: Kim, Taehyeon, et al.
Published: (2024)
Guiding Reasoning in Small Language Models with LLM Assistance
by: Kim, Yujin, et al.
Published: (2025)
by: Kim, Yujin, et al.
Published: (2025)
A Jointly Efficient and Optimal Algorithm for Heteroskedastic Generalized Linear Bandits with Adversarial Corruptions
by: Kim, Sanghwa, et al.
Published: (2026)
by: Kim, Sanghwa, et al.
Published: (2026)
Multi-Drafter Speculative Decoding with Alignment Feedback
by: Kim, Taehyeon, et al.
Published: (2026)
by: Kim, Taehyeon, et al.
Published: (2026)
MERIT Feedback Elicits Better Bargaining in LLM Negotiators
by: Oh, Jihwan, et al.
Published: (2026)
by: Oh, Jihwan, et al.
Published: (2026)
MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning Segmentation
by: Jang, Donggon, et al.
Published: (2025)
by: Jang, Donggon, et al.
Published: (2025)
PROGrasp: Pragmatic Human-Robot Communication for Object Grasping
by: Kang, Gi-Cheon, et al.
Published: (2023)
by: Kang, Gi-Cheon, et al.
Published: (2023)
Instance-Optimal Estimation with Multiple LLM Judges on a Budget
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
by: Jang, Kangwook, et al.
Published: (2023)
by: Jang, Kangwook, et al.
Published: (2023)
Predicting LLM Reasoning Performance with Small Proxy Model
by: Koh, Woosung, et al.
Published: (2025)
by: Koh, Woosung, et al.
Published: (2025)
Generative Visual Code Mobile World Models
by: Koh, Woosung, et al.
Published: (2026)
by: Koh, Woosung, et al.
Published: (2026)
Influencing optimistic bias: Moderating roles of perceived severity and proximity
by: Hyuksoo Kim, et al.
Published: (2024)
by: Hyuksoo Kim, et al.
Published: (2024)
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
by: Chang, Fu-Chieh, et al.
Published: (2024)
by: Chang, Fu-Chieh, et al.
Published: (2024)
V-STaR: Training Verifiers for Self-Taught Reasoners
by: Hosseini, Arian, et al.
Published: (2024)
by: Hosseini, Arian, et al.
Published: (2024)
STaR-SQL: Self-Taught Reasoner for Text-to-SQL
by: He, Mingqian, et al.
Published: (2025)
by: He, Mingqian, et al.
Published: (2025)
Diffusion-based Episodes Augmentation for Offline Multi-Agent Reinforcement Learning
by: Oh, Jihwan, et al.
Published: (2024)
by: Oh, Jihwan, et al.
Published: (2024)
Chondroprotective Effect of Extract in Primary Chondrocytes and Rat OA Model.
by: Jang, Ji Yun, et al.
Published: (2024)
by: Jang, Ji Yun, et al.
Published: (2024)
Soft Inductive Bias Approach via Explicit Reasoning Perspectives in Inappropriate Utterance Detection Using Large Language Models
by: Kim, Ju-Young, et al.
Published: (2025)
by: Kim, Ju-Young, et al.
Published: (2025)
On Dirac equations with Hartree type nonlinearity in modulation spaces
by: Kim, Seongyeon, et al.
Published: (2023)
by: Kim, Seongyeon, et al.
Published: (2023)
Flooding with Absorption: An Efficient Protocol for Heterogeneous Bandits over Complex Networks
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
DeClotH: Decomposable 3D Cloth and Human Body Reconstruction from a Single Image
by: Nam, Hyeongjin, et al.
Published: (2025)
by: Nam, Hyeongjin, et al.
Published: (2025)
Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters
by: Yi, Euiin, et al.
Published: (2024)
by: Yi, Euiin, et al.
Published: (2024)
Hypernetwork-Driven Model Fusion for Federated Domain Generalization
by: Bartholet, Marc, et al.
Published: (2024)
by: Bartholet, Marc, et al.
Published: (2024)
Three‐Dimensional Efficient Myelin‐Weighted Imaging Utilizing Direct Visualization of Short Transverse Relaxation Time Component (ViSTa)
by: Se‐Hong Oh, et al.
Published: (2025)
by: Se‐Hong Oh, et al.
Published: (2025)
FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models
by: Lee, Seunghan, et al.
Published: (2026)
by: Lee, Seunghan, et al.
Published: (2026)
HS-STaR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation
by: Xiong, Feng, et al.
Published: (2025)
by: Xiong, Feng, et al.
Published: (2025)
Regular Schur labeled skew shape posets and their 0-Hecke modules
by: Kim, Young-Hun, et al.
Published: (2023)
by: Kim, Young-Hun, et al.
Published: (2023)
Re3val: Reinforced and Reranked Generative Retrieval
by: Song, EuiYul, et al.
Published: (2024)
by: Song, EuiYul, et al.
Published: (2024)
FedSOL: Stabilized Orthogonal Learning with Proximal Restrictions in Federated Learning
by: Lee, Gihun, et al.
Published: (2023)
by: Lee, Gihun, et al.
Published: (2023)
A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
by: Kim, SuYeon, et al.
Published: (2026)
by: Kim, SuYeon, et al.
Published: (2026)
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
DBMovi-GS: Dynamic View Synthesis from Blurry Monocular Video via Sparse-Controlled Gaussian Splatting
by: Song, Yeon-Ji, et al.
Published: (2025)
by: Song, Yeon-Ji, et al.
Published: (2025)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
by: Kim, Sungnyun, et al.
Published: (2024)
by: Kim, Sungnyun, et al.
Published: (2024)
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
by: Kim, Hayeon, et al.
Published: (2025)
by: Kim, Hayeon, et al.
Published: (2025)
A Unified Confidence Sequence for Generalized Linear Models, with Applications to Bandits
by: Lee, Junghyun, et al.
Published: (2024)
by: Lee, Junghyun, et al.
Published: (2024)
Improved Regret Bounds of (Multinomial) Logistic Bandits via Regret-to-Confidence-Set Conversion
by: Lee, Junghyun, et al.
Published: (2023)
by: Lee, Junghyun, et al.
Published: (2023)
Integrability as an attractor of adiabatic flows
by: Kim, Hyeongjin, et al.
Published: (2023)
by: Kim, Hyeongjin, et al.
Published: (2023)
Similar Items
-
FlickerFusion: Intra-trajectory Domain Generalizing Multi-Agent RL
by: Koh, Woosung, et al.
Published: (2024) -
Instructive Decoding: Instruction-Tuned Large Language Models are Self-Refiner from Noisy Instructions
by: Kim, Taehyeon, et al.
Published: (2023) -
$C^2$: Scalable Auto-Feedback for LLM-based Chart Generation
by: Koh, Woosung, et al.
Published: (2024) -
Revisiting Early-Learning Regularization When Federated Learning Meets Noisy Labels
by: Kim, Taehyeon, et al.
Published: (2024) -
Guiding Reasoning in Small Language Models with LLM Assistance
by: Kim, Yujin, et al.
Published: (2025)