Preference Distillation via Value based Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kwon, Minchan, Ko, Junwon, Kim, Kangil, Kim, Junmo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models
by: Kwon, Minchan, et al.
Published: (2024)
by: Kwon, Minchan, et al.
Published: (2024)
SFLD: Reducing the content bias for AI-generated Image Detection
by: Gye, Seoyeon, et al.
Published: (2025)
by: Gye, Seoyeon, et al.
Published: (2025)
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023)
by: Kim, Hoyong, et al.
Published: (2023)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction
by: Park, Jinwook, et al.
Published: (2025)
by: Park, Jinwook, et al.
Published: (2025)
Structural Optimization Ambiguity and Simplicity Bias in Unsupervised Neural Grammar Induction
by: Park, Jinwook, et al.
Published: (2024)
by: Park, Jinwook, et al.
Published: (2024)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024)
by: Choi, Jaehyun, et al.
Published: (2024)
Feature Structure Distillation with Centered Kernel Alignment in BERT Transferring
by: Jung, Hee-Jun, et al.
Published: (2022)
by: Jung, Hee-Jun, et al.
Published: (2022)
ConceptPrism: Concept Disentanglement in Personalized Diffusion Models via Residual Token Optimization
by: Kim, Minseo, et al.
Published: (2026)
by: Kim, Minseo, et al.
Published: (2026)
RSCF: Relation-Semantics Consistent Filter for Entity Embedding of Knowledge Graph
by: Kim, Junsik, et al.
Published: (2025)
by: Kim, Junsik, et al.
Published: (2025)
Comparison Reveals Commonality: Customized Image Generation through Contrastive Inversion
by: Kim, Minseo, et al.
Published: (2025)
by: Kim, Minseo, et al.
Published: (2025)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
by: Kim, Jongsuk, et al.
Published: (2025)
by: Kim, Jongsuk, et al.
Published: (2025)
FxSearcher: gradient-free text-driven audio transformation
by: Ki, Hojoon, et al.
Published: (2025)
by: Ki, Hojoon, et al.
Published: (2025)
Bayesian Multi-Task Transfer Learning for Soft Prompt Tuning
by: Lee, Haeju, et al.
Published: (2024)
by: Lee, Haeju, et al.
Published: (2024)
BAPO: Base-Anchored Preference Optimization for Overcoming Forgetting in Large Language Models Personalization
by: Lee, Gihun, et al.
Published: (2024)
by: Lee, Gihun, et al.
Published: (2024)
MedRep: Medical Concept Representation for General Electronic Health Record Foundation Models
by: Kim, Junmo, et al.
Published: (2025)
by: Kim, Junmo, et al.
Published: (2025)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
by: Kim, Jongsuk, et al.
Published: (2024)
by: Kim, Jongsuk, et al.
Published: (2024)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)
by: Ko, Jongwoo, et al.
Published: (2026)
From Belief Entrenchment to Robust Reasoning in LLM Agents
by: Oh, Jihwan, et al.
Published: (2025)
by: Oh, Jihwan, et al.
Published: (2025)
Instruct-4DGS: Efficient Dynamic Scene Editing via 4D Gaussian-based Static-Dynamic Separation
by: Kwon, Joohyun, et al.
Published: (2025)
by: Kwon, Joohyun, et al.
Published: (2025)
Do Music Preferences Reflect Cultural Values? A Cross-National Analysis Using Music Embedding and World Values Survey
by: Kim, Yongjae, et al.
Published: (2025)
by: Kim, Yongjae, et al.
Published: (2025)
Dialogue Systems for Emotional Support via Value Reinforcement
by: Kim, Juhee, et al.
Published: (2025)
by: Kim, Juhee, et al.
Published: (2025)
Forget What Matters, Keep the Rest: Selective Unlearning of Informative Tokens
by: Koh, Seunghee, et al.
Published: (2026)
by: Koh, Seunghee, et al.
Published: (2026)
RLKD: Distilling LLMs' Reasoning via Reinforcement Learning
by: Xu, Shicheng, et al.
Published: (2025)
by: Xu, Shicheng, et al.
Published: (2025)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
by: Kim, Minwu, et al.
Published: (2025)
by: Kim, Minwu, et al.
Published: (2025)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
by: Kim, Minchan, et al.
Published: (2024)
by: Kim, Minchan, et al.
Published: (2024)
Stable-GFlowNet: Toward Diverse and Robust LLM Red-Teaming via Contrastive Trajectory Balance
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Subgraph-Aware Training of Language Models for Knowledge Graph Completion Using Structure-Aware Contrastive Learning
by: Ko, Youmin, et al.
Published: (2024)
by: Ko, Youmin, et al.
Published: (2024)
Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
by: Bai, Fengshuo, et al.
Published: (2024)
by: Bai, Fengshuo, et al.
Published: (2024)
Dependency Parsing with the Structuralized Prompt Template
by: Kim, Keunha, et al.
Published: (2025)
by: Kim, Keunha, et al.
Published: (2025)
DistiLLM: Towards Streamlined Distillation for Large Language Models
by: Ko, Jongwoo, et al.
Published: (2024)
by: Ko, Jongwoo, et al.
Published: (2024)
Aligning to Thousands of Preferences via System Message Generalization
by: Lee, Seongyun, et al.
Published: (2024)
by: Lee, Seongyun, et al.
Published: (2024)
General Preference Reinforcement Learning
by: Umer, Muhammad, et al.
Published: (2026)
by: Umer, Muhammad, et al.
Published: (2026)
CTPD: Cross Tokenizer Preference Distillation
by: Nguyen, Truong, et al.
Published: (2026)
by: Nguyen, Truong, et al.
Published: (2026)
Reverse Engineering Human Preferences with Reinforcement Learning
by: Alazraki, Lisa, et al.
Published: (2025)
by: Alazraki, Lisa, et al.
Published: (2025)
PLaD: Preference-based Large Language Model Distillation with Pseudo-Preference Pairs
by: Zhang, Rongzhi, et al.
Published: (2024)
by: Zhang, Rongzhi, et al.
Published: (2024)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Separating Novel Features for Logical Anomaly Detection: A Straightforward yet Effective Approach
by: Lee, Kangil, et al.
Published: (2024)
by: Lee, Kangil, et al.
Published: (2024)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
by: Shen, Yiyang, et al.
Published: (2026)
by: Shen, Yiyang, et al.
Published: (2026)
Similar Items
-
StablePrompt: Automatic Prompt Tuning using Reinforcement Learning for Large Language Models
by: Kwon, Minchan, et al.
Published: (2024) -
SFLD: Reducing the content bias for AI-generated Image Detection
by: Gye, Seoyeon, et al.
Published: (2025) -
Revisiting Softmax Masking: Stop Gradient for Enhancing Stability in Replay-based Continual Learning
by: Kim, Hoyong, et al.
Published: (2023) -
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026) -
Probability Distribution Collapse: A Critical Bottleneck to Compact Unsupervised Neural Grammar Induction
by: Park, Jinwook, et al.
Published: (2025)