Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Liang, Ren-Wei, Hsu, Chin-Ting, Yu, Chan-Hung, Agrawal, Saransh, Huang, Shih-Cheng, Lin, Chieh-Yen, Chen, Shang-Tse, Huang, Kuan-Hao, Sun, Shao-Hua |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SHA256 at SemEval-2025 Task 4: Selective Amnesia -- Constrained Unlearning for Large Language Models via Knowledge Isolation
par: Agrawal, Saransh, et autres
Publié: (2025)
par: Agrawal, Saransh, et autres
Publié: (2025)
Dishonesty in Helpful and Harmless Alignment
par: Huang, Youcheng, et autres
Publié: (2024)
par: Huang, Youcheng, et autres
Publié: (2024)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
par: Huang, Yue, et autres
Publié: (2025)
par: Huang, Yue, et autres
Publié: (2025)
A$^2$TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation
par: Hsu, Sheng-Chi, et autres
Publié: (2026)
par: Hsu, Sheng-Chi, et autres
Publié: (2026)
VDR-LLM-Prolog: Alignment: Helpful, Harmless, Honest Through Structure, Not Interference
par: Howland, Geoffrey
Publié: (2026)
par: Howland, Geoffrey
Publié: (2026)
H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
par: Tekin, Selim Furkan, et autres
Publié: (2024)
par: Tekin, Selim Furkan, et autres
Publié: (2024)
Contrastive Bi-Projector for Unsupervised Domain Adaption
par: Huang, Lin-Chieh, et autres
Publié: (2023)
par: Huang, Lin-Chieh, et autres
Publié: (2023)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
par: Chehbouni, Khaoula, et autres
Publié: (2024)
par: Chehbouni, Khaoula, et autres
Publié: (2024)
Evaluating the Potential Immunostimulatory Effects of Cryptomeria japonica Leaf Essential Oil on Honey Bees (Apis mellifera)
par: Hao‐Yung Wang, et autres
Publié: (2025)
par: Hao‐Yung Wang, et autres
Publié: (2025)
Lymphocyte‐related ratios in methamphetamine‐induced psychotic disorder in Taiwan, comparing with patients with schizophrenia
par: Mei‐Hing Ng, et autres
Publié: (2024)
par: Mei‐Hing Ng, et autres
Publié: (2024)
PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head
par: Lin, Chieh-Yen, et autres
Publié: (2026)
par: Lin, Chieh-Yen, et autres
Publié: (2026)
Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages
par: Huang, Shih-Cheng, et autres
Publié: (2023)
par: Huang, Shih-Cheng, et autres
Publié: (2023)
Too Helpful, Too Harmless, Too Honest or Just Right?
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
Revisiting Semi-supervised Adversarial Robustness via Noise-aware Online Robust Distillation
par: Wu, Tsung-Han, et autres
Publié: (2024)
par: Wu, Tsung-Han, et autres
Publié: (2024)
Towards Accurate Heart Rate Measurement from Ultra-Short Video Clips via Periodicity-Guided rPPG Estimation and Signal Reconstruction
par: Huanga, Pei-Kai, et autres
Publié: (2025)
par: Huanga, Pei-Kai, et autres
Publié: (2025)
Distribution Shift Alignment Helps LLMs Simulate Survey Response Distributions
par: Huang, Ji, et autres
Publié: (2025)
par: Huang, Ji, et autres
Publié: (2025)
Towards Harmless Multimodal Assistants with Blind Preference Optimization
par: Li, Yongqi, et autres
Publié: (2025)
par: Li, Yongqi, et autres
Publié: (2025)
Host–Guest Complexation of α‐Cyclodextrin and Triiodide Ions for Enhanced Performance of Ionic Thermoelectric Capacitors
par: Shih‐Ting Kao, et autres
Publié: (2025)
par: Shih‐Ting Kao, et autres
Publié: (2025)
Host–Guest Complexation of α‐Cyclodextrin and Triiodide Ions for Enhanced Performance of Ionic Thermoelectric Capacitors (Adv. Energy Mater. 15/2025)
par: Shih‐Ting Kao, et autres
Publié: (2025)
par: Shih‐Ting Kao, et autres
Publié: (2025)
Harmless Yet Harmful: Neutral Prompting Attacks for Stealthy Hallucination Steering in Agent Skills
par: Hsu, Chia-Yi, et autres
Publié: (2026)
par: Hsu, Chia-Yi, et autres
Publié: (2026)
HRLAIF: Improvements in Helpfulness and Harmlessness in Open-domain Reinforcement Learning From AI Feedback
par: Li, Ang, et autres
Publié: (2024)
par: Li, Ang, et autres
Publié: (2024)
InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance
par: Wang, Pengyu, et autres
Publié: (2024)
par: Wang, Pengyu, et autres
Publié: (2024)
Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
par: Liu, Max, et autres
Publié: (2024)
par: Liu, Max, et autres
Publié: (2024)
Dosimetric evaluation and clinical application of collimated apertures with proton beam line scanning in stereotactic radiotherapy
par: Chen‐Yu Chou, et autres
Publié: (2025)
par: Chen‐Yu Chou, et autres
Publié: (2025)
Warp-STAR: High-performance, Differentiable GPU-Accelerated Static Timing Analysis through Warp-oriented Parallel Orchestration
par: Huang, En-Ming, et autres
Publié: (2026)
par: Huang, En-Ming, et autres
Publié: (2026)
Endoscopic Excision of Transsellar Transsphenoidal Meningoencephalocele Utilizing the Slip‐Knot Technique
par: Chin‐Nung Liu, et autres
Publié: (2025)
par: Chin‐Nung Liu, et autres
Publié: (2025)
SciCapenter: Supporting Caption Composition for Scientific Figures with Machine-Generated Captions and Ratings
par: Hsu, Ting-Yao, et autres
Publié: (2024)
par: Hsu, Ting-Yao, et autres
Publié: (2024)
Calo-VQ: Vector-Quantized Two-Stage Generative Model in Calorimeter Simulation
par: Liu, Qibin, et autres
Publié: (2024)
par: Liu, Qibin, et autres
Publié: (2024)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
par: Le, Linh, et autres
Publié: (2026)
par: Le, Linh, et autres
Publié: (2026)
Visually Descriptive Language Model for Vector Graphics Reasoning
par: Wang, Zhenhailong, et autres
Publié: (2024)
par: Wang, Zhenhailong, et autres
Publié: (2024)
Comment on ‘Beyond the scoreboard: Coaches’ UV‐related skin cancer knowledge in outdoor sports’
par: Luo‐Wei Chan, et autres
Publié: (2025)
par: Luo‐Wei Chan, et autres
Publié: (2025)
A Constructor-Theoretic and Quantum Information Approach to the Three-Step Photoemission Model: A Theoretical Investigation
par: Malhotra, Saransh
Publié: (2025)
par: Malhotra, Saransh
Publié: (2025)
How Does Conversation Length Impact User's Satisfaction? A Case Study of Length-Controlled Conversations with LLM-Powered Chatbots
par: Huang, Shih-Hong, et autres
Publié: (2024)
par: Huang, Shih-Hong, et autres
Publié: (2024)
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence
par: Kim, Minbeom, et autres
Publié: (2024)
par: Kim, Minbeom, et autres
Publié: (2024)
Adaptive Unknown Fault Detection and Few-Shot Continual Learning for Condition Monitoring in Ultrasonic Metal Welding
par: Eslaminia, Ahmadreza, et autres
Publié: (2026)
par: Eslaminia, Ahmadreza, et autres
Publié: (2026)
RPRO: Ranked Preference Reinforcement Optimization for Enhancing Medical QA and Diagnostic Reasoning
par: Hsu, Chia-Hsuan, et autres
Publié: (2025)
par: Hsu, Chia-Hsuan, et autres
Publié: (2025)
Effective Manipulation of Water Droplets on Open Superhydrophobic Glass Surfaces by Using a Triboelectrically Charged Polytetrafluoroethylene Rod on the Back Side of These Surfaces
par: Wei Chen Huang, et autres
Publié: (2025)
par: Wei Chen Huang, et autres
Publié: (2025)
Effective Manipulation of Water Droplets on Open Superhydrophobic Glass Surfaces by Using a Triboelectrically Charged Polytetrafluoroethylene Rod on the Back Side of These Surfaces (Adv. Mater. Interfaces 12/2025)
par: Wei Chen Huang, et autres
Publié: (2025)
par: Wei Chen Huang, et autres
Publié: (2025)
Central sleep apnea as an initial presentation of small cell lung carcinoma with anti‐Hu antibody‐related paraneoplastic neurologic syndrome
par: Yi‐Tse Su, et autres
Publié: (2024)
par: Yi‐Tse Su, et autres
Publié: (2024)
Steering Vector Fields for Context-Aware Inference-Time Control in Large Language Models
par: Li, Jiaqian, et autres
Publié: (2026)
par: Li, Jiaqian, et autres
Publié: (2026)
Documents similaires
-
SHA256 at SemEval-2025 Task 4: Selective Amnesia -- Constrained Unlearning for Large Language Models via Knowledge Isolation
par: Agrawal, Saransh, et autres
Publié: (2025) -
Dishonesty in Helpful and Harmless Alignment
par: Huang, Youcheng, et autres
Publié: (2024) -
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
par: Huang, Yue, et autres
Publié: (2025) -
A$^2$TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation
par: Hsu, Sheng-Chi, et autres
Publié: (2026) -
VDR-LLM-Prolog: Alignment: Helpful, Harmless, Honest Through Structure, Not Interference
par: Howland, Geoffrey
Publié: (2026)