Aligning Language Models Using Follow-up Likelihood as Reward Signal
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Chen, Chong, Dading, Jiang, Feng, Tang, Chengguang, Gao, Anningzhe, Tang, Guohua, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
von: Zhang, Chen, et al.
Veröffentlicht: (2024)
Unsupervised Mutual Learning of Discourse Parsing and Topic Segmentation in Dialogue
von: Xu, Jiahui, et al.
Veröffentlicht: (2024)
von: Xu, Jiahui, et al.
Veröffentlicht: (2024)
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
von: Jiang, Songtao, et al.
Veröffentlicht: (2025)
Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
von: Yu, Zichao, et al.
Veröffentlicht: (2025)
von: Yu, Zichao, et al.
Veröffentlicht: (2025)
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2023)
von: Yu, Fei, et al.
Veröffentlicht: (2023)
GraphWiz: An Instruction-Following Language Model for Graph Problems
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
von: Chen, Nuo, et al.
Veröffentlicht: (2024)
Outcome Accuracy is Not Enough: Aligning the Reasoning Process of Reward Models
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
von: Wang, Binghai, et al.
Veröffentlicht: (2026)
Efficient Tuning and Inference for Large Language Models on Textual Graphs
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
von: Zhu, Yun, et al.
Veröffentlicht: (2024)
IHEval: Evaluating Language Models on Following the Instruction Hierarchy
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihan, et al.
Veröffentlicht: (2025)
Enabling Doctor-Centric Medical AI with LLMs through Workflow-Aligned Tasks and Benchmarks
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
von: Xie, Wenya, et al.
Veröffentlicht: (2025)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
von: Zhang, Qingru, et al.
Veröffentlicht: (2025)
Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
von: Zhou, Li, et al.
Veröffentlicht: (2025)
von: Zhou, Li, et al.
Veröffentlicht: (2025)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
Atoxia: Red-teaming Large Language Models with Target Toxic Answers
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
von: Du, Yuhao, et al.
Veröffentlicht: (2024)
LLMs for Mathematical Modeling: Towards Bridging the Gap between Natural and Mathematical Languages
von: Huang, Xuhan, et al.
Veröffentlicht: (2024)
von: Huang, Xuhan, et al.
Veröffentlicht: (2024)
Add-One-In: Incremental Sample Selection for Large Language Models via a Choice-Based Greedy Paradigm
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
von: Li, Zhuo, et al.
Veröffentlicht: (2025)
Checklists Are Better Than Reward Models For Aligning Language Models
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
von: Viswanathan, Vijay, et al.
Veröffentlicht: (2025)
Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
von: Ma, Qiyao, et al.
Veröffentlicht: (2026)
Prior Constraints-based Reward Model Training for Aligning Large Language Models
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
von: Zhou, Hang, et al.
Veröffentlicht: (2024)
ViLBench: A Suite for Vision-Language Process Reward Modeling
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
von: Tu, Haoqin, et al.
Veröffentlicht: (2025)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
von: Jiang, Shuyu, et al.
Veröffentlicht: (2023)
Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language Models
von: Liu, Ziche, et al.
Veröffentlicht: (2024)
von: Liu, Ziche, et al.
Veröffentlicht: (2024)
When Reward Hacking Rebounds: Understanding and Mitigating It with Representation-Level Signals
von: Wu, Rui, et al.
Veröffentlicht: (2026)
von: Wu, Rui, et al.
Veröffentlicht: (2026)
SR-GRPO: Stable Rank as an Intrinsic Geometric Reward for Large Language Model Alignment
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
von: Tang, Yixuan, et al.
Veröffentlicht: (2025)
Bridging Research and Readers: A Multi-Modal Automated Academic Papers Interpretation System
von: Jiang, Feng, et al.
Veröffentlicht: (2024)
von: Jiang, Feng, et al.
Veröffentlicht: (2024)
Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text
von: Yang, Lingyi, et al.
Veröffentlicht: (2023)
von: Yang, Lingyi, et al.
Veröffentlicht: (2023)
Empirical Study of Mutual Reinforcement Effect and Application in Few-shot Text Classification Tasks via Prompt
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
Unleashing the Potentials of Likelihood Composition for Multi-modal Language Models
von: Zhao, Shitian, et al.
Veröffentlicht: (2024)
von: Zhao, Shitian, et al.
Veröffentlicht: (2024)
Domain Mixture Design via Log-Likelihood Differences for Aligning Language Models with a Target Model
von: Kishino, Ryo, et al.
Veröffentlicht: (2026)
von: Kishino, Ryo, et al.
Veröffentlicht: (2026)
Fin-PRM: A Domain-Specialized Process Reward Model for Financial Reasoning in Large Language Models
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
von: Zhu, Jie, et al.
Veröffentlicht: (2025)
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
von: Li, Xiaomin, et al.
Veröffentlicht: (2025)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
von: Wu, Keming, et al.
Veröffentlicht: (2025)
von: Wu, Keming, et al.
Veröffentlicht: (2025)
Application of LLM Agents in Recruitment: A Novel Framework for Resume Screening
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
von: Gan, Chengguang, et al.
Veröffentlicht: (2024)
URPO: A Unified Reward & Policy Optimization Framework for Large Language Models
von: Lu, Songshuo, et al.
Veröffentlicht: (2025)
von: Lu, Songshuo, et al.
Veröffentlicht: (2025)
Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
von: Ji, Jiaming, et al.
Veröffentlicht: (2024)
Aligning Language Models with Real-time Knowledge Editing
von: Tang, Chenming, et al.
Veröffentlicht: (2025)
von: Tang, Chenming, et al.
Veröffentlicht: (2025)
Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of Translationese
von: Liu, Yikang, et al.
Veröffentlicht: (2025)
von: Liu, Yikang, et al.
Veröffentlicht: (2025)
KARMA: Karma-Aligned Reward Model Adaptation
von: Scott, Jared, et al.
Veröffentlicht: (2026)
von: Scott, Jared, et al.
Veröffentlicht: (2026)
MemoryRewardBench: Benchmarking Reward Models for Long-Term Memory Management in Large Language Models
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
von: Tang, Zecheng, et al.
Veröffentlicht: (2026)
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
von: Si, Shuzheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TS-Align: A Teacher-Student Collaborative Framework for Scalable Iterative Finetuning of Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2024) -
Unsupervised Mutual Learning of Discourse Parsing and Topic Segmentation in Dialogue
von: Xu, Jiahui, et al.
Veröffentlicht: (2024) -
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
von: Jiang, Songtao, et al.
Veröffentlicht: (2025) -
Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
von: Yu, Zichao, et al.
Veröffentlicht: (2025) -
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2023)