Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tuan, Yi-Lin, Chen, Xilun, Smith, Eric Michael, Martin, Louis, Batra, Soumya, Celikyilmaz, Asli, Wang, William Yang, Bikel, Daniel M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Auditing LLM Benchmarks with Item Response Theory
von: Land, Sander, et al.
Veröffentlicht: (2026)
von: Land, Sander, et al.
Veröffentlicht: (2026)
Backtracking Improves Generation Safety
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
von: Tan, Yingshui, et al.
Veröffentlicht: (2025)
von: Tan, Yingshui, et al.
Veröffentlicht: (2025)
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
Open-Domain Text Evaluation via Contrastive Distribution Methods
von: Lu, Sidi, et al.
Veröffentlicht: (2023)
von: Lu, Sidi, et al.
Veröffentlicht: (2023)
Learning to Interrupt in Language-based Multi-agent Communication
von: Wang, Danqing, et al.
Veröffentlicht: (2026)
von: Wang, Danqing, et al.
Veröffentlicht: (2026)
Bi-Factorial Preference Optimization: Balancing Safety-Helpfulness in Language Models
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Wenxuan, et al.
Veröffentlicht: (2024)
Adaptive Decoding via Latent Preference Optimization
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
von: Dhuliawala, Shehzaad, et al.
Veröffentlicht: (2024)
Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
von: Zhao, Zhenyu, et al.
Veröffentlicht: (2026)
Don't throw away your value model! Generating more preferable text with Value-Guided Monte-Carlo Tree Search decoding
von: Liu, Jiacheng, et al.
Veröffentlicht: (2023)
von: Liu, Jiacheng, et al.
Veröffentlicht: (2023)
The Majority is not always right: RL training for solution aggregation
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
von: Zhao, Wenting, et al.
Veröffentlicht: (2025)
reWordBench: Benchmarking and Improving the Robustness of Reward Models with Transformed Inputs
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2025)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2025)
EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
von: Li, Yuxuan, et al.
Veröffentlicht: (2025)
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
von: Shen, Guobin, et al.
Veröffentlicht: (2024)
Resprompt: Residual Connection Prompting Advances Multi-Step Reasoning in Large Language Models
von: Jiang, Song, et al.
Veröffentlicht: (2023)
von: Jiang, Song, et al.
Veröffentlicht: (2023)
A Gradient Analysis Framework for Rewarding Good and Penalizing Bad Examples in Language Models
von: Tuan, Yi-Lin, et al.
Veröffentlicht: (2024)
von: Tuan, Yi-Lin, et al.
Veröffentlicht: (2024)
Libra-Leaderboard: Towards Responsible AI through a Balanced Leaderboard of Safety and Capability
von: Li, Haonan, et al.
Veröffentlicht: (2024)
von: Li, Haonan, et al.
Veröffentlicht: (2024)
DRAMA: Diverse Augmentation from Large Language Models to Smaller Dense Retrievers
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
von: Ma, Xueguang, et al.
Veröffentlicht: (2025)
FLAME: Factuality-Aware Alignment for Large Language Models
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
von: Lin, Sheng-Chieh, et al.
Veröffentlicht: (2024)
Acceleration from a Phase of Entropic Balance
von: Chakrabarti, Soumya
Veröffentlicht: (2025)
von: Chakrabarti, Soumya
Veröffentlicht: (2025)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
von: Yang, Jinluan, et al.
Veröffentlicht: (2025)
Paying Less Generalization Tax: A Cross-Domain Generalization Study of RL Training for LLM Agents
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
von: Liu, Zhihan, et al.
Veröffentlicht: (2026)
Balancing Safety and Helpfulness in Healthcare AI Assistants through Iterative Preference Alignment
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
von: Nghiem, Huy, et al.
Veröffentlicht: (2025)
HonestLLM: Toward an Honest and Helpful Large Language Model
von: Gao, Chujie, et al.
Veröffentlicht: (2024)
von: Gao, Chujie, et al.
Veröffentlicht: (2024)
Cold-Start Personalization via Training-Free Priors from Structured World Models
von: Bose, Avinandan, et al.
Veröffentlicht: (2026)
von: Bose, Avinandan, et al.
Veröffentlicht: (2026)
Improving Faithfulness of Abstractive Summarization by Controlling Confounding Effect of Irrelevant Sentences
von: Ghoshal, Asish, et al.
Veröffentlicht: (2022)
von: Ghoshal, Asish, et al.
Veröffentlicht: (2022)
Towards Comprehensive Post Safety Alignment of Large Language Models via Safety Patching
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2024)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
von: Sclar, Melanie, et al.
Veröffentlicht: (2024)
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
von: Guo, William, et al.
Veröffentlicht: (2025)
von: Guo, William, et al.
Veröffentlicht: (2025)
MedicalBench: Evaluating Large Language Models Toward Improved Medical Concept Extraction
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
von: Yang, Zhichao, et al.
Veröffentlicht: (2026)
Meta-Contrastive Learning for Vision-Language Models via Task-Adaptive CLIP Training
von: Fouladvand, Merham, et al.
Veröffentlicht: (2026)
von: Fouladvand, Merham, et al.
Veröffentlicht: (2026)
Hermitian-Einstein Metrics on Parabolic Bundles over compact complex surfaces
von: Li, Xilun, et al.
Veröffentlicht: (2025)
von: Li, Xilun, et al.
Veröffentlicht: (2025)
New exotic examples of Ricci limit spaces
von: Li, Xilun, et al.
Veröffentlicht: (2024)
von: Li, Xilun, et al.
Veröffentlicht: (2024)
On the shrinking solitons of generalized Ricci flow
von: Li, Xilun, et al.
Veröffentlicht: (2024)
von: Li, Xilun, et al.
Veröffentlicht: (2024)
Diversity Helps Jailbreak Large Language Models
von: Zhao, Weiliang, et al.
Veröffentlicht: (2024)
von: Zhao, Weiliang, et al.
Veröffentlicht: (2024)
One Trigger Token Is Enough: A Defense Strategy for Balancing Safety and Usability in Large Language Models
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
von: Gu, Haoran, et al.
Veröffentlicht: (2025)
Accurate Failure Prediction in Agents Does Not Imply Effective Failure Prevention
von: Vasudev, Rakshith, et al.
Veröffentlicht: (2026)
von: Vasudev, Rakshith, et al.
Veröffentlicht: (2026)
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework
von: Qi, Yupeng, et al.
Veröffentlicht: (2025)
von: Qi, Yupeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Auditing LLM Benchmarks with Item Response Theory
von: Land, Sander, et al.
Veröffentlicht: (2026) -
Backtracking Improves Generation Safety
von: Zhang, Yiming, et al.
Veröffentlicht: (2024) -
Equilibrate RLHF: Towards Balancing Helpfulness-Safety Trade-off in Large Language Models
von: Tan, Yingshui, et al.
Veröffentlicht: (2025) -
Post-training an LLM for RAG? Train on Self-Generated Demonstrations
von: Finlayson, Matthew, et al.
Veröffentlicht: (2025) -
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
von: Saha, Swarnadeep, et al.
Veröffentlicht: (2023)