Sample-Efficient Alignment for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zichen, Chen, Changyu, Du, Chao, Lee, Wee Sun, Lin, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Understanding R1-Zero-Like Training: A Critical Perspective
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Locality Sensitive Sparse Encoding for Learning World Models Online
von: Liu, Zichen, et al.
Veröffentlicht: (2024)
von: Liu, Zichen, et al.
Veröffentlicht: (2024)
GEM: A Gym for Agentic LLMs
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Continual Reinforcement Learning by Planning with Online World Models
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
von: Liu, Zichen, et al.
Veröffentlicht: (2025)
Language Models Can Learn from Verbal Feedback Without Scalar Rewards
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
von: Luo, Renjie, et al.
Veröffentlicht: (2025)
SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Variational Reasoning for Language Models
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
Towards Scalable Automated Alignment of LLMs: A Survey
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
von: Cao, Boxi, et al.
Veröffentlicht: (2024)
Aligners: Decoupling LLMs and Alignment
von: Ngweta, Lilian, et al.
Veröffentlicht: (2024)
von: Ngweta, Lilian, et al.
Veröffentlicht: (2024)
Enhancing Time Series Forecasting via Multi-Level Text Alignment with LLMs
von: Zhao, Taibiao, et al.
Veröffentlicht: (2025)
von: Zhao, Taibiao, et al.
Veröffentlicht: (2025)
Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
von: Dou, Longxu, et al.
Veröffentlicht: (2025)
von: Dou, Longxu, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
Lifelong Safety Alignment for Language Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
von: Chen, Hong, et al.
Veröffentlicht: (2024)
von: Chen, Hong, et al.
Veröffentlicht: (2024)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
von: Li, Guanlin, et al.
Veröffentlicht: (2025)
von: Li, Guanlin, et al.
Veröffentlicht: (2025)
Towards a copilot in BIM authoring tool using a large language model-based agent for intelligent human-machine interaction
von: Du, Changyu, et al.
Veröffentlicht: (2024)
von: Du, Changyu, et al.
Veröffentlicht: (2024)
A Closer Look at Machine Unlearning for Large Language Models
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
von: Hu, Haiquan, et al.
Veröffentlicht: (2025)
The Alignment Tax: Response Homogenization in Aligned LLMs and Its Implications for Uncertainty Estimation
von: Liu, Mingyi
Veröffentlicht: (2026)
von: Liu, Mingyi
Veröffentlicht: (2026)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
Direct Alignment of Draft Model for Speculative Decoding with Chat-Fine-Tuned LLMs
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
von: Goel, Raghavv, et al.
Veröffentlicht: (2024)
When Attention Sink Emerges in Language Models: An Empirical View
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
von: Gu, Xiangming, et al.
Veröffentlicht: (2024)
Efficiently Distilling LLMs for Edge Applications
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
von: Kundu, Achintya, et al.
Veröffentlicht: (2024)
Pluralistic Alignment for Healthcare: A Role-Driven Framework
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
von: Zhong, Jiayou, et al.
Veröffentlicht: (2025)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
von: Kim, Dongyoung, et al.
Veröffentlicht: (2024)
Aligner: Efficient Alignment by Learning to Correct
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
von: Azmat, Muneeza, et al.
Veröffentlicht: (2025)
von: Azmat, Muneeza, et al.
Veröffentlicht: (2025)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
von: Patel, Dev, et al.
Veröffentlicht: (2025)
von: Patel, Dev, et al.
Veröffentlicht: (2025)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
HAL: Inducing Human-likeness in LLMs with Alignment
von: Hasan, Masum, et al.
Veröffentlicht: (2026)
von: Hasan, Masum, et al.
Veröffentlicht: (2026)
Reward-free Alignment for Conflicting Objectives
von: Chen, Peter, et al.
Veröffentlicht: (2026)
von: Chen, Peter, et al.
Veröffentlicht: (2026)
Replay Failures as Successes: Sample-Efficient Reinforcement Learning for Instruction Following
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
von: Zhang, Kongcheng, et al.
Veröffentlicht: (2025)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
von: Li, Changhao, et al.
Veröffentlicht: (2024)
von: Li, Changhao, et al.
Veröffentlicht: (2024)
Value Augmented Sampling for Language Model Alignment and Personalization
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
von: Han, Seungwook, et al.
Veröffentlicht: (2024)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
von: Wang, Fei, et al.
Veröffentlicht: (2025)
von: Wang, Fei, et al.
Veröffentlicht: (2025)
Efficient multi-prompt evaluation of LLMs
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
von: Polo, Felipe Maia, et al.
Veröffentlicht: (2024)
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
von: Du, Hongzhe, et al.
Veröffentlicht: (2025)
Bootstrapping Language Models with DPO Implicit Rewards
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
von: Chen, Changyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Understanding R1-Zero-Like Training: A Critical Perspective
von: Liu, Zichen, et al.
Veröffentlicht: (2025) -
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025) -
Defeating the Training-Inference Mismatch via FP16
von: Qi, Penghui, et al.
Veröffentlicht: (2025) -
Locality Sensitive Sparse Encoding for Learning World Models Online
von: Liu, Zichen, et al.
Veröffentlicht: (2024) -
GEM: A Gym for Agentic LLMs
von: Liu, Zichen, et al.
Veröffentlicht: (2025)