Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
Fuente:
arXiv
Salvato in:
| Autori principali: | Dang, Quy-Anh, Ngo, Chris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024)
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
di: Gong, Shuzhi, et al.
Pubblicazione: (2026)
di: Gong, Shuzhi, et al.
Pubblicazione: (2026)
MoD: A Distribution-Based Approach for Merging Large Language Models
di: Dang, Quy-Anh, et al.
Pubblicazione: (2024)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2024)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)
di: Nayak, Nihal V., et al.
Pubblicazione: (2026)
di: Nayak, Nihal V., et al.
Pubblicazione: (2026)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
di: Cooper, A. Feder, et al.
Pubblicazione: (2024)
RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2025)
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study
di: Lv, Keyu, et al.
Pubblicazione: (2026)
di: Lv, Keyu, et al.
Pubblicazione: (2026)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
di: He, Qianxi, et al.
Pubblicazione: (2025)
di: He, Qianxi, et al.
Pubblicazione: (2025)
Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
di: You, Kisung
Pubblicazione: (2025)
di: You, Kisung
Pubblicazione: (2025)
Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
di: Clark, Tyler, et al.
Pubblicazione: (2025)
di: Clark, Tyler, et al.
Pubblicazione: (2025)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
di: Yu, Zony, et al.
Pubblicazione: (2025)
di: Yu, Zony, et al.
Pubblicazione: (2025)
What do Transformers Know about Government?
di: Hou, Jue, et al.
Pubblicazione: (2024)
di: Hou, Jue, et al.
Pubblicazione: (2024)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
Library Designs Revisited: What Works--What Doesn't.
di: Metz, T. John, et al.
Pubblicazione: (1987)
di: Metz, T. John, et al.
Pubblicazione: (1987)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
Explorations of the Softmax Space: Knowing When the Neural Network Doesn't Know
di: Sikar, Daniel, et al.
Pubblicazione: (2025)
di: Sikar, Daniel, et al.
Pubblicazione: (2025)
What Happens When Small Is Made Smaller? Exploring the Impact of Compression on Small Data Pretrained Language Models
di: Awobade, Busayo, et al.
Pubblicazione: (2024)
di: Awobade, Busayo, et al.
Pubblicazione: (2024)
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
di: Sernau, Luke
Pubblicazione: (2024)
di: Sernau, Luke
Pubblicazione: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
di: Deng, Wenhao, et al.
Pubblicazione: (2025)
di: Deng, Wenhao, et al.
Pubblicazione: (2025)
Small LLMs Do Not Learn a Generalizable Theory of Mind via Reinforcement Learning
di: Sarangi, Sneheel, et al.
Pubblicazione: (2025)
di: Sarangi, Sneheel, et al.
Pubblicazione: (2025)
When More Data Doesn't Help: Limits of Adaptation in Multitask Learning
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
di: Hanneke, Steve, et al.
Pubblicazione: (2026)
wav2graph: A Framework for Supervised Learning Knowledge Graph from Speech
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
di: Le-Duc, Khai, et al.
Pubblicazione: (2024)
Measuring What LLMs Think They Do: SHAP Faithfulness and Deployability on Financial Tabular Classification
di: AlMarri, Saeed, et al.
Pubblicazione: (2025)
di: AlMarri, Saeed, et al.
Pubblicazione: (2025)
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
di: Zheng, Chujie, et al.
Pubblicazione: (2025)
di: Zheng, Chujie, et al.
Pubblicazione: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
di: Zverev, Egor, et al.
Pubblicazione: (2024)
di: Zverev, Egor, et al.
Pubblicazione: (2024)
Table-r1: Self-supervised and Reinforcement Learning for Program-based Table Reasoning in Small Language Models
di: Jin, Rihui, et al.
Pubblicazione: (2025)
di: Jin, Rihui, et al.
Pubblicazione: (2025)
TRN-R1-Zero: Text-rich Network Reasoning via LLMs with Reinforcement Learning Only
di: Liu, Yilun, et al.
Pubblicazione: (2026)
di: Liu, Yilun, et al.
Pubblicazione: (2026)
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs
di: Lin, Jiacheng, et al.
Pubblicazione: (2025)
di: Lin, Jiacheng, et al.
Pubblicazione: (2025)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
di: He, Di, et al.
Pubblicazione: (2026)
di: He, Di, et al.
Pubblicazione: (2026)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
di: He, Xuan, et al.
Pubblicazione: (2024)
di: He, Xuan, et al.
Pubblicazione: (2024)
Reasoning Models Don't Always Say What They Think
di: Chen, Yanda, et al.
Pubblicazione: (2025)
di: Chen, Yanda, et al.
Pubblicazione: (2025)
Learning to Reason in LLMs by Expectation Maximization
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
di: Lee, Junghyun, et al.
Pubblicazione: (2025)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
di: Suvarna, Ashima, et al.
Pubblicazione: (2026)
di: Suvarna, Ashima, et al.
Pubblicazione: (2026)
What Do Language Models Hear? Probing for Auditory Representations in Language Models
di: Ngo, Jerry, et al.
Pubblicazione: (2024)
di: Ngo, Jerry, et al.
Pubblicazione: (2024)
What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
di: Li, Ming, et al.
Pubblicazione: (2025)
di: Li, Ming, et al.
Pubblicazione: (2025)
MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
di: Krohn-Grimberghe, Artus
Pubblicazione: (2026)
di: Krohn-Grimberghe, Artus
Pubblicazione: (2026)
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
di: Damirchi, Hamed, et al.
Pubblicazione: (2026)
di: Damirchi, Hamed, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
di: Berlot-Attwell, Ian, et al.
Pubblicazione: (2024) -
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026) -
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
di: Gong, Shuzhi, et al.
Pubblicazione: (2026) -
MoD: A Distribution-Based Approach for Merging Large Language Models
di: Dang, Quy-Anh, et al.
Pubblicazione: (2024) -
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)