Salvato in:
| Autori principali: | Zheng, Chen, Sun, Ke, Wu, Hang, Xi, Chenguang, Zhou, Xun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2403.02513 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs
di: Zheng, Chen, et al.
Pubblicazione: (2024)
di: Zheng, Chen, et al.
Pubblicazione: (2024)
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers
di: Zheng, Chen, et al.
Pubblicazione: (2024)
di: Zheng, Chen, et al.
Pubblicazione: (2024)
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
di: Zheng, Chen, et al.
Pubblicazione: (2025)
di: Zheng, Chen, et al.
Pubblicazione: (2025)
WPO: Enhancing RLHF with Weighted Preference Optimization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024)
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
di: Kong, Jiawei, et al.
Pubblicazione: (2025)
di: Kong, Jiawei, et al.
Pubblicazione: (2025)
Dishonesty in Helpful and Harmless Alignment
di: Huang, Youcheng, et al.
Pubblicazione: (2024)
di: Huang, Youcheng, et al.
Pubblicazione: (2024)
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
di: Zhou, Qinhao, et al.
Pubblicazione: (2024)
FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
di: Cui, Shiyao, et al.
Pubblicazione: (2023)
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond
di: Zheng, Shen, et al.
Pubblicazione: (2023)
di: Zheng, Shen, et al.
Pubblicazione: (2023)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
di: Yang, Jinluan, et al.
Pubblicazione: (2025)
di: Yang, Jinluan, et al.
Pubblicazione: (2025)
Harnessing RLHF for Robust Unanswerability Recognition and Trustworthy Response Generation in LLMs
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
di: Lin, Shuyuan, et al.
Pubblicazione: (2025)
Taming Overconfidence in LLMs: Reward Calibration in RLHF
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
di: Leng, Jixuan, et al.
Pubblicazione: (2024)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
di: Wang, Pengyu, et al.
Pubblicazione: (2024)
di: Wang, Pengyu, et al.
Pubblicazione: (2024)
Towards Federated RLHF with Aggregated Client Preference for LLMs
di: Wu, Feijie, et al.
Pubblicazione: (2024)
di: Wu, Feijie, et al.
Pubblicazione: (2024)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
di: Tekin, Selim Furkan, et al.
Pubblicazione: (2024)
di: Tekin, Selim Furkan, et al.
Pubblicazione: (2024)
A Desideratum for Conversational Agents: Capabilities, Challenges, and Future Directions
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
di: Acikgoz, Emre Can, et al.
Pubblicazione: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
di: Ji, Jiaming, et al.
Pubblicazione: (2024)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
di: Xi, Zhiheng, et al.
Pubblicazione: (2025)
di: Xi, Zhiheng, et al.
Pubblicazione: (2025)
Continual SFT Matches Multimodal RLHF with Negative Supervision
di: Zhu, Ke, et al.
Pubblicazione: (2024)
di: Zhu, Ke, et al.
Pubblicazione: (2024)
AAPO: Enhancing the Reasoning Capabilities of LLMs with Advantage Margin
di: Xiong, Jian, et al.
Pubblicazione: (2025)
di: Xiong, Jian, et al.
Pubblicazione: (2025)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
di: Chehbouni, Khaoula, et al.
Pubblicazione: (2024)
MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions
di: Chai, Yekun, et al.
Pubblicazione: (2024)
di: Chai, Yekun, et al.
Pubblicazione: (2024)
DuplexMamba: Enhancing Real-time Speech Conversations with Duplex and Streaming Capabilities
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
di: Lu, Xiangyu, et al.
Pubblicazione: (2025)
More Than Catastrophic Forgetting: Integrating General Capabilities For Domain-Specific LLMs
di: Liu, Chengyuan, et al.
Pubblicazione: (2024)
di: Liu, Chengyuan, et al.
Pubblicazione: (2024)
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities
di: Sun, Chung-En, et al.
Pubblicazione: (2024)
di: Sun, Chung-En, et al.
Pubblicazione: (2024)
MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities
di: Chen, Jingxue, et al.
Pubblicazione: (2025)
di: Chen, Jingxue, et al.
Pubblicazione: (2025)
RLHF Workflow: From Reward Modeling to Online RLHF
di: Dong, Hanze, et al.
Pubblicazione: (2024)
di: Dong, Hanze, et al.
Pubblicazione: (2024)
OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
di: Hu, Jian, et al.
Pubblicazione: (2024)
di: Hu, Jian, et al.
Pubblicazione: (2024)
Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
di: Jørgenvåg, Magnus, et al.
Pubblicazione: (2026)
di: Jørgenvåg, Magnus, et al.
Pubblicazione: (2026)
Too Helpful, Too Harmless, Too Honest or Just Right?
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
di: Kashyap, Gautam Siddharth, et al.
Pubblicazione: (2025)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
di: Xu, Jing, et al.
Pubblicazione: (2026)
di: Xu, Jing, et al.
Pubblicazione: (2026)
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
di: Zhou, Ziya, et al.
Pubblicazione: (2024)
di: Zhou, Ziya, et al.
Pubblicazione: (2024)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
di: Li, Yongqi, et al.
Pubblicazione: (2025)
di: Li, Yongqi, et al.
Pubblicazione: (2025)
Segmenting Text and Learning Their Rewards for Improved RLHF in Language Model
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
di: Yin, Yueqin, et al.
Pubblicazione: (2025)
Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
di: Spiegelhalter, Urs, et al.
Pubblicazione: (2025)
di: Spiegelhalter, Urs, et al.
Pubblicazione: (2025)
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
di: Wang, Xiaobo, et al.
Pubblicazione: (2026)
di: Wang, Xiaobo, et al.
Pubblicazione: (2026)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
di: Chen, Ruxiao, et al.
Pubblicazione: (2025)
di: Chen, Ruxiao, et al.
Pubblicazione: (2025)
Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs
di: Zheng, Chen, et al.
Pubblicazione: (2024) -
ICE-GRT: Instruction Context Enhancement by Generative Reinforcement based Transformers
di: Zheng, Chen, et al.
Pubblicazione: (2024) -
Balanced Actor Initialization: Stable RLHF Training of Distillation-Based Reasoning Models
di: Zheng, Chen, et al.
Pubblicazione: (2025) -
WPO: Enhancing RLHF with Weighted Preference Optimization
di: Zhou, Wenxuan, et al.
Pubblicazione: (2024) -
Revisiting Backdoor Attacks on LLMs: A Stealthy and Practical Poisoning Framework via Harmless Inputs
di: Kong, Jiawei, et al.
Pubblicazione: (2025)