FIRST: Teach A Reliable Large Language Model Through Efficient Trustworthy Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Shum, KaShun, Xu, Minrui, Zhang, Jianshu, Chen, Zixin, Diao, Shizhe, Dong, Hanze, Zhang, Jipeng, Raza, Muhammad Omer |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
by: Shum, KaShun, et al.
Published: (2023)
by: Shum, KaShun, et al.
Published: (2023)
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
by: Diao, Shizhe, et al.
Published: (2023)
by: Diao, Shizhe, et al.
Published: (2023)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
SWE-RM: Execution-free Feedback For Software Engineering Agents
by: Shum, KaShun, et al.
Published: (2025)
by: Shum, KaShun, et al.
Published: (2025)
Plum: Prompt Learning using Metaheuristic
by: Pan, Rui, et al.
Published: (2023)
by: Pan, Rui, et al.
Published: (2023)
MLLM-Protector: Ensuring MLLM's Safety without Hurting Performance
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Entropy-Regularized Process Reward Model
by: Zhang, Hanning, et al.
Published: (2024)
by: Zhang, Hanning, et al.
Published: (2024)
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
by: Wang, Ruida, et al.
Published: (2024)
by: Wang, Ruida, et al.
Published: (2024)
Active Prompting with Chain-of-Thought for Large Language Models
by: Diao, Shizhe, et al.
Published: (2023)
by: Diao, Shizhe, et al.
Published: (2023)
Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoning
by: Yang, Zhicheng, et al.
Published: (2026)
by: Yang, Zhicheng, et al.
Published: (2026)
The Instinctive Bias: Spurious Images lead to Illusion in MLLMs
by: Han, Tianyang, et al.
Published: (2024)
by: Han, Tianyang, et al.
Published: (2024)
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
by: Xiong, Zixin, et al.
Published: (2026)
by: Xiong, Zixin, et al.
Published: (2026)
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
by: Liu, Mingjie, et al.
Published: (2025)
by: Liu, Mingjie, et al.
Published: (2025)
ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge Bases
by: Do, Quyet V., et al.
Published: (2024)
by: Do, Quyet V., et al.
Published: (2024)
MA-LoT: Model-Collaboration Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
MegaFlow: Large-Scale Distributed Orchestration System for the Agentic Era
by: Zhang, Lei, et al.
Published: (2026)
by: Zhang, Lei, et al.
Published: (2026)
$\textbf{PLUM}$: Improving Code LMs with Execution-Guided On-Policy Preference Learning Driven By Synthetic Test Cases
by: Zhang, Dylan, et al.
Published: (2024)
by: Zhang, Dylan, et al.
Published: (2024)
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
by: Pang, Bo, et al.
Published: (2025)
by: Pang, Bo, et al.
Published: (2025)
Towards Trustworthy Dataset Distillation
by: Ma, Shijie, et al.
Published: (2023)
by: Ma, Shijie, et al.
Published: (2023)
Personalized Visual Instruction Tuning
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Mitigating the Alignment Tax of RLHF
by: Lin, Yong, et al.
Published: (2023)
by: Lin, Yong, et al.
Published: (2023)
Bridge-Coder: Unlocking LLMs' Potential to Overcome Language Gaps in Low-Resource Code
by: Zhang, Jipeng, et al.
Published: (2024)
by: Zhang, Jipeng, et al.
Published: (2024)
STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
by: Guo, Yichen, et al.
Published: (2025)
by: Guo, Yichen, et al.
Published: (2025)
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
by: Zhao, Zirui, et al.
Published: (2024)
by: Zhao, Zirui, et al.
Published: (2024)
RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
by: Niu, Cheng, et al.
Published: (2023)
by: Niu, Cheng, et al.
Published: (2023)
R-Tuning: Instructing Large Language Models to Say `I Don't Know'
by: Zhang, Hanning, et al.
Published: (2023)
by: Zhang, Hanning, et al.
Published: (2023)
SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales
by: Xu, Tianyang, et al.
Published: (2024)
by: Xu, Tianyang, et al.
Published: (2024)
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
by: Chen, Zixin, et al.
Published: (2026)
by: Chen, Zixin, et al.
Published: (2026)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
Thrust: Adaptively Propels Large Language Models with External Knowledge
by: Zhao, Xinran, et al.
Published: (2023)
by: Zhao, Xinran, et al.
Published: (2023)
Uncertainty-Aware and Decoder-Aligned Learning for Video Summarization
by: Tariq, Omer, et al.
Published: (2026)
by: Tariq, Omer, et al.
Published: (2026)
MMIDR: Teaching Large Language Model to Interpret Multimodal Misinformation via Knowledge Distillation
by: Wang, Longzheng, et al.
Published: (2024)
by: Wang, Longzheng, et al.
Published: (2024)
The Challenge of Teaching Reasoning to LLMs Without RL or Distillation
by: Du, Wei, et al.
Published: (2025)
by: Du, Wei, et al.
Published: (2025)
RNR: Teaching Large Language Models to Follow Roles and Rules
by: Wang, Kuan, et al.
Published: (2024)
by: Wang, Kuan, et al.
Published: (2024)
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
by: Wang, Ruida, et al.
Published: (2025)
by: Wang, Ruida, et al.
Published: (2025)
A Reliability Theory of Compromise Decisions for Large-Scale Stochastic Programs
by: Diao, Shuotao, et al.
Published: (2024)
by: Diao, Shuotao, et al.
Published: (2024)
Similar Items
-
Automatic Prompt Augmentation and Selection with Chain-of-Thought from Labeled Data
by: Shum, KaShun, et al.
Published: (2023) -
LMFlow: An Extensible Toolkit for Finetuning and Inference of Large Foundation Models
by: Diao, Shizhe, et al.
Published: (2023) -
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024) -
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
by: Pan, Rui, et al.
Published: (2025) -
SWE-RM: Execution-free Feedback For Software Engineering Agents
by: Shum, KaShun, et al.
Published: (2025)