Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Pareja, Aldo, Nayak, Nikhil Shivakumar, Wang, Hao, Killamsetty, Krishnateja, Sudalairaj, Shivchander, Zhao, Wenlong, Han, Seungwook, Bhandwaldar, Abhishek, Xu, Guangxuan, Xu, Kai, Han, Ligong, Inglis, Luke, Srivastava, Akash |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
by: Nayak, Nikhil Shivakumar, et al.
Published: (2025)
by: Nayak, Nikhil Shivakumar, et al.
Published: (2025)
Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning
by: Chanda, Prateek, et al.
Published: (2025)
by: Chanda, Prateek, et al.
Published: (2025)
Hopscotch: Discovering and Skipping Redundancies in Language Models
by: Eyceoz, Mustafa, et al.
Published: (2025)
by: Eyceoz, Mustafa, et al.
Published: (2025)
Graph Attention for Heterogeneous Graphs with Positional Encoding
by: Nayak, Nikhil Shivakumar
Published: (2025)
by: Nayak, Nikhil Shivakumar
Published: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
by: Yang, Shan
Published: (2026)
by: Yang, Shan
Published: (2026)
Multi-Paradigm Agent Interaction in Practice:A Systematic Analysis of Generator-Evaluator, ReAct Loop,and Adversarial Evaluation in the buddyMe Framework
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
Active Context Compression: Autonomous Memory Management in LLM Agents
by: Verma, Nikhil
Published: (2026)
by: Verma, Nikhil
Published: (2026)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
by: Kashyap, Pankhi, et al.
Published: (2024)
by: Kashyap, Pankhi, et al.
Published: (2024)
Ultra-Reduced-Impact-Encased-Logging (URIEL): propose a new method for selective sustainable logging and post-harvest silvicultural treatment in tropical forest using airborne robotics systems
by: Albiero, Daniel, et al.
Published: (2026)
by: Albiero, Daniel, et al.
Published: (2026)
TrafficRAG: A Multimodal RAG Framework for Traffic Accident Liability Determination
by: Li, Xu, et al.
Published: (2026)
by: Li, Xu, et al.
Published: (2026)
5G Traffic Prediction with Time Series Analysis
by: Nayak, Nikhil, et al.
Published: (2021)
by: Nayak, Nikhil, et al.
Published: (2021)
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026)
by: Atreja, Dhruv, et al.
Published: (2026)
An Explainable Collaborative Dialogue System using a Theory of Mind
by: Cohen, Philip R., et al.
Published: (2023)
by: Cohen, Philip R., et al.
Published: (2023)
Predicting and improving test-time scaling laws via reward tail-guided search
by: Li, Muheng, et al.
Published: (2026)
by: Li, Muheng, et al.
Published: (2026)
Good to Go: The LOOP Skill Engine That Hits 99% Success and Slashes Token Usage by 99% via One-Shot Recording and Deterministic Replay
by: Wang, Xiaohua, et al.
Published: (2026)
by: Wang, Xiaohua, et al.
Published: (2026)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Factored Diffusion Policies:Compositionally Generalized Robot Control with a Single Score Network
by: Mitra, Sayan, et al.
Published: (2026)
by: Mitra, Sayan, et al.
Published: (2026)
Deployment-Time Reliability of Learned Robot Policies
by: Agia, Christopher
Published: (2026)
by: Agia, Christopher
Published: (2026)
Gyan: An Explainable Neuro-Symbolic Language Model
by: Srinivasan, Venkat, et al.
Published: (2026)
by: Srinivasan, Venkat, et al.
Published: (2026)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
by: Sharma, Anika, et al.
Published: (2025)
by: Sharma, Anika, et al.
Published: (2025)
Meta-Evaluation of Translation Evaluation Methods: a systematic up-to-date overview
by: Han, Lifeng, et al.
Published: (2016)
by: Han, Lifeng, et al.
Published: (2016)
GraphWalk: Enabling Reasoning in Large Language Models through Tool-Based Graph Navigation
by: Ghandi, Taraneh, et al.
Published: (2026)
by: Ghandi, Taraneh, et al.
Published: (2026)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
by: Tang, Wenjie, et al.
Published: (2026)
by: Tang, Wenjie, et al.
Published: (2026)
A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19
by: Khandelwal, Vedant, et al.
Published: (2024)
by: Khandelwal, Vedant, et al.
Published: (2024)
N-Agent Ad Hoc Teamwork
by: Wang, Caroline, et al.
Published: (2024)
by: Wang, Caroline, et al.
Published: (2024)
Automated Circuit Interpretation via Probe Prompting
by: Birardi, Giuseppe
Published: (2025)
by: Birardi, Giuseppe
Published: (2025)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
by: Pareschi, Remo
Published: (2025)
by: Pareschi, Remo
Published: (2025)
A Super-Learner with Large Language Models for Medical Emergency Advising
by: Aityan, Sergey K., et al.
Published: (2025)
by: Aityan, Sergey K., et al.
Published: (2025)
When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judges
by: Darshan, Parth, et al.
Published: (2026)
by: Darshan, Parth, et al.
Published: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
by: Gupta, Sunny, et al.
Published: (2024)
by: Gupta, Sunny, et al.
Published: (2024)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
by: Fang, Qitong, et al.
Published: (2026)
by: Fang, Qitong, et al.
Published: (2026)
IMUVIE: Pickup Timeline Action Localization via Motion Movies
by: Clapham, John, et al.
Published: (2024)
by: Clapham, John, et al.
Published: (2024)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
by: Chen, Junting, et al.
Published: (2024)
by: Chen, Junting, et al.
Published: (2024)
EvoCUA: Evolving Computer Use Agents via Learning from Scalable Synthetic Experience
by: Xue, Taofeng, et al.
Published: (2026)
by: Xue, Taofeng, et al.
Published: (2026)
Can AI Assist in Olympiad Coding
by: Ren, Samuel
Published: (2025)
by: Ren, Samuel
Published: (2025)
Multi-Agent Design Assistant for the Simulation of Inertial Fusion Energy
by: Shachar, Meir H., et al.
Published: (2025)
by: Shachar, Meir H., et al.
Published: (2025)
OmniAcc: Personalized Accessibility Assistant Using Generative AI
by: Karki, Siddhant, et al.
Published: (2025)
by: Karki, Siddhant, et al.
Published: (2025)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
by: Römer, Ralf, et al.
Published: (2025)
by: Römer, Ralf, et al.
Published: (2025)
Similar Items
-
Sculpting Subspaces: Constrained Full Fine-Tuning in LLMs for Continual Learning
by: Nayak, Nikhil Shivakumar, et al.
Published: (2025) -
Learning What Matters: Probabilistic Task Selection via Mutual Information for Model Finetuning
by: Chanda, Prateek, et al.
Published: (2025) -
Hopscotch: Discovering and Skipping Redundancies in Language Models
by: Eyceoz, Mustafa, et al.
Published: (2025) -
Graph Attention for Heterogeneous Graphs with Positional Encoding
by: Nayak, Nikhil Shivakumar
Published: (2025) -
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
by: Yang, Shan
Published: (2026)