Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Singh, Navan Preet, Wang, Xiaokun, Garikipati, Anurag, Ciobanu, Madalina, Mao, Qingqing, Das, Ritankar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
von: Chen, Shaolong, et al.
Veröffentlicht: (2026)
Narrowing the Gap: Supervised Fine-Tuning of Open-Source LLMs as a Viable Alternative to Proprietary Models for Pedagogical Tools
von: Solano, Lorenzo Lee, et al.
Veröffentlicht: (2025)
von: Solano, Lorenzo Lee, et al.
Veröffentlicht: (2025)
Towards Pedagogical LLMs with Supervised Fine Tuning for Computing Education
von: Vassar, Alexandra, et al.
Veröffentlicht: (2024)
von: Vassar, Alexandra, et al.
Veröffentlicht: (2024)
Supervised Fine-Tuning LLMs to Behave as Pedagogical Agents in Programming Education
von: Ross, Emily, et al.
Veröffentlicht: (2025)
von: Ross, Emily, et al.
Veröffentlicht: (2025)
GraphGen: Enhancing Supervised Fine-Tuning for LLMs with Knowledge-Driven Synthetic Data Generation
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
von: Chen, Zihong, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Constrained Optimization for Cloud Application Deployment via Graph Neural Networks and Satisfiability Modulo Theory
von: Erascu, Madalina
Veröffentlicht: (2025)
von: Erascu, Madalina
Veröffentlicht: (2025)
Explaining Fine Tuned LLMs via Counterfactuals A Knowledge Graph Driven Framework
von: Wang, Yucheng, et al.
Veröffentlicht: (2025)
von: Wang, Yucheng, et al.
Veröffentlicht: (2025)
Study on Benthic organisms of Ghazal Ozan River in Zanjan Province
von: Navan Maghsoodi, M.
Veröffentlicht: (2013)
von: Navan Maghsoodi, M.
Veröffentlicht: (2013)
SDA: Steering-Driven Distribution Alignment for Open LLMs without Fine-Tuning
von: Xia, Wei, et al.
Veröffentlicht: (2025)
von: Xia, Wei, et al.
Veröffentlicht: (2025)
UltraLink: An Open-Source Knowledge-Enhanced Multilingual Supervised Fine-tuning Dataset
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024)
von: Sun, Hao
Veröffentlicht: (2024)
UFT: Unifying Supervised and Reinforcement Fine-Tuning
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
von: Liu, Mingyang, et al.
Veröffentlicht: (2025)
NEFMind: Parameter-Efficient Fine-Tuning of Open-Source LLMs for Telecom APIs Automation
von: Khan, Zainab, et al.
Veröffentlicht: (2025)
von: Khan, Zainab, et al.
Veröffentlicht: (2025)
Dos casos de calambre refractario del escribano en la clínica de dolor: ¿está la respuesta en la toxina botulínica?
von: Preet Mohinder Singh
Veröffentlicht: (2013)
von: Preet Mohinder Singh
Veröffentlicht: (2013)
Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning
von: Alizadeh, Meysam, et al.
Veröffentlicht: (2023)
von: Alizadeh, Meysam, et al.
Veröffentlicht: (2023)
Thomason's completion for K-theory and cyclic homology of quotient stacks
von: Krishna, Amalendu, et al.
Veröffentlicht: (2025)
von: Krishna, Amalendu, et al.
Veröffentlicht: (2025)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Pedagogically-Inspired Data Synthesis for Language Model Knowledge Distillation
von: He, Bowei, et al.
Veröffentlicht: (2026)
von: He, Bowei, et al.
Veröffentlicht: (2026)
Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning
von: Mecklenburg, Nick, et al.
Veröffentlicht: (2024)
von: Mecklenburg, Nick, et al.
Veröffentlicht: (2024)
Efficient Differentially Private Fine-Tuning of LLMs via Reinforcement Learning
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
von: Khadangi, Afshin, et al.
Veröffentlicht: (2025)
Excite, Attend and Segment (EASe): Domain-Agnostic Fine-Grained Mask Discovery with Feature Calibration and Self-Supervised Upsampling
von: Singh, Deepank, et al.
Veröffentlicht: (2026)
von: Singh, Deepank, et al.
Veröffentlicht: (2026)
The inference of Fokker-Planck equations via transport maps
von: Han, Saem, et al.
Veröffentlicht: (2025)
von: Han, Saem, et al.
Veröffentlicht: (2025)
SFT-GO: Supervised Fine-Tuning with Group Optimization for Large Language Models
von: Kim, Gyuhak, et al.
Veröffentlicht: (2025)
von: Kim, Gyuhak, et al.
Veröffentlicht: (2025)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
Learning to Staff: Offline Reinforcement Learning and Fine-Tuned LLMs for Warehouse Staffing Optimization
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2026)
von: Kujanpää, Kalle, et al.
Veröffentlicht: (2026)
Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs
von: Ovadia, Oded, et al.
Veröffentlicht: (2023)
von: Ovadia, Oded, et al.
Veröffentlicht: (2023)
Improving 2D Feature Representations by 3D-Aware Fine-Tuning
von: Yue, Yuanwen, et al.
Veröffentlicht: (2024)
von: Yue, Yuanwen, et al.
Veröffentlicht: (2024)
Optimizing Recommendations using Fine-Tuned LLMs
von: Cheema, Prabhdeep, et al.
Veröffentlicht: (2025)
von: Cheema, Prabhdeep, et al.
Veröffentlicht: (2025)
Rethinking Scale: The Efficacy of Fine-Tuned Open-Source LLMs in Large-Scale Reproducible Social Science Research
von: Carammia, Marcello, et al.
Veröffentlicht: (2024)
von: Carammia, Marcello, et al.
Veröffentlicht: (2024)
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
von: Qu, Yuxiao, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
von: Baroian, Andrei
Veröffentlicht: (2025)
von: Baroian, Andrei
Veröffentlicht: (2025)
Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
von: Rallapalli, Swati, et al.
Veröffentlicht: (2025)
von: Rallapalli, Swati, et al.
Veröffentlicht: (2025)
KG-FIT: Knowledge Graph Fine-Tuning Upon Open-World Knowledge
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
von: Jiang, Pengcheng, et al.
Veröffentlicht: (2024)
Supervised Fine Tuning on Curated Data is Reinforcement Learning (and can be improved)
von: Qin, Chongli, et al.
Veröffentlicht: (2025)
von: Qin, Chongli, et al.
Veröffentlicht: (2025)
Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning
von: Tang, Zhenchao, et al.
Veröffentlicht: (2025)
von: Tang, Zhenchao, et al.
Veröffentlicht: (2025)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
von: Gekhman, Zorik, et al.
Veröffentlicht: (2024)
Interplay of Exchange and Anisotropy Energies in Landscape of Current‐Driven Skyrmion in Defect‐Engineered Racetracks
von: Preet Kamal, et al.
Veröffentlicht: (2026)
von: Preet Kamal, et al.
Veröffentlicht: (2026)
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
von: Wang, Jingyao, et al.
Veröffentlicht: (2025)
von: Wang, Jingyao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
State-of-the-Art Arabic Language Modeling with Sparse MoE Fine-Tuning and Chain-of-Thought Distillation
von: Singh, Navan Preet, et al.
Veröffentlicht: (2026) -
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024) -
AdaDPO: Self-Adaptive Direct Preference Optimization with Balanced Gradient Updates
von: Chen, Shaolong, et al.
Veröffentlicht: (2026) -
Narrowing the Gap: Supervised Fine-Tuning of Open-Source LLMs as a Viable Alternative to Proprietary Models for Pedagogical Tools
von: Solano, Lorenzo Lee, et al.
Veröffentlicht: (2025) -
Towards Pedagogical LLMs with Supervised Fine Tuning for Computing Education
von: Vassar, Alexandra, et al.
Veröffentlicht: (2024)