H3Fusion: Helpful, Harmless, Honest Fusion of Aligned LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Tekin, Selim Furkan, Ilhan, Fatih, Huang, Tiansheng, Hu, Sihao, Xu, Yichang, Yahn, Zachary, Liu, Ling |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025)
by: Huang, Tiansheng, et al.
Published: (2025)
A Survey on Large Language Model-Based Game Agents
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity
by: Tekin, Selim Furkan, et al.
Published: (2024)
by: Tekin, Selim Furkan, et al.
Published: (2024)
Adversarial Attention Perturbations for Large Object Detection Transformers
by: Yahn, Zachary, et al.
Published: (2025)
by: Yahn, Zachary, et al.
Published: (2025)
Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Personalized Face Privacy Protection From a Single Image
by: Yahn, Zachary, et al.
Published: (2026)
by: Yahn, Zachary, et al.
Published: (2026)
A Multi-Agent Perception-Action Alliance for Efficient Long Video Reasoning
by: Xu, Yichang, et al.
Published: (2026)
by: Xu, Yichang, et al.
Published: (2026)
Lisa: Lazy Safety Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient Decoding
by: Ilhan, Fatih, et al.
Published: (2026)
by: Ilhan, Fatih, et al.
Published: (2026)
A Neurosymbolic Agent System for Compositional Visual Reasoning
by: Xu, Yichang, et al.
Published: (2025)
by: Xu, Yichang, et al.
Published: (2025)
MELT: A Behavioral Trace Dataset for High-Risk Memecoin Launch Detection
by: Hu, Sihao, et al.
Published: (2026)
by: Hu, Sihao, et al.
Published: (2026)
PokeLLMon: A Human-Parity Agent for Pokemon Battles with Large Language Models
by: Hu, Sihao, et al.
Published: (2024)
by: Hu, Sihao, et al.
Published: (2024)
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
by: Tekin, Selim Furkan, et al.
Published: (2025)
by: Tekin, Selim Furkan, et al.
Published: (2025)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Large Language Model based Smart Contract Auditing with LLMBugScanner
by: Yuan, Yining, et al.
Published: (2025)
by: Yuan, Yining, et al.
Published: (2025)
Robust Few-Shot Ensemble Learning with Focal Diversity-Based Pruning
by: Tekin, Selim Furkan, et al.
Published: (2024)
by: Tekin, Selim Furkan, et al.
Published: (2024)
Too Helpful, Too Harmless, Too Honest or Just Right?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
HonestLLM: Toward an Honest and Helpful Large Language Model
by: Gao, Chujie, et al.
Published: (2024)
by: Gao, Chujie, et al.
Published: (2024)
Vaccine: Perturbation-aware Alignment for Large Language Models against Harmful Fine-tuning Attack
by: Huang, Tiansheng, et al.
Published: (2024)
by: Huang, Tiansheng, et al.
Published: (2024)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Servant, Stalker, Predator: How An Honest, Helpful, And Harmless (3H) Agent Unlocks Adversarial Skills
by: Noever, David
Published: (2025)
by: Noever, David
Published: (2025)
Dishonesty in Helpful and Harmless Alignment
by: Huang, Youcheng, et al.
Published: (2024)
by: Huang, Youcheng, et al.
Published: (2024)
B+ANN: A Fast Billion-Scale Disk-based Nearest-Neighbor Index
by: Tekin, Selim Furkan, et al.
Published: (2025)
by: Tekin, Selim Furkan, et al.
Published: (2025)
FedHFT: Efficient Federated Finetuning with Heterogeneous Edge Clients
by: Ilhan, Fatih, et al.
Published: (2025)
by: Ilhan, Fatih, et al.
Published: (2025)
Fusion-Eval: Integrating Assistant Evaluators with LLMs
by: Shu, Lei, et al.
Published: (2023)
by: Shu, Lei, et al.
Published: (2023)
BeHonest: Benchmarking Honesty in Large Language Models
by: Chern, Steffi, et al.
Published: (2024)
by: Chern, Steffi, et al.
Published: (2024)
SportsMetrics: Blending Text and Numerical Data to Understand Information Fusion in LLMs
by: Hu, Yebowen, et al.
Published: (2024)
by: Hu, Yebowen, et al.
Published: (2024)
Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging
by: Yang, Jinluan, et al.
Published: (2025)
by: Yang, Jinluan, et al.
Published: (2025)
Prioritization First, Principles Second: An Adaptive Interpretation of Helpful, Honest, and Harmless Principles
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Adaptive Helpfulness-Harmlessness Alignment with Preference Vectors
by: Liang, Ren-Wei, et al.
Published: (2025)
by: Liang, Ren-Wei, et al.
Published: (2025)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion
by: Wu, Yutong, et al.
Published: (2025)
by: Wu, Yutong, et al.
Published: (2025)
Beyond the Safety Bundle: Auditing the Helpful and Harmless Dataset
by: Chehbouni, Khaoula, et al.
Published: (2024)
by: Chehbouni, Khaoula, et al.
Published: (2024)
The Moltbook Files: A Harmless Slopocalypse or Humanity's Last Experiment
by: Brach, William, et al.
Published: (2026)
by: Brach, William, et al.
Published: (2026)
VDR-LLM-Prolog: Alignment: Helpful, Harmless, Honest Through Structure, Not Interference
by: Howland, Geoffrey
Published: (2026)
by: Howland, Geoffrey
Published: (2026)
ReFusion: Improving Natural Language Understanding with Computation-Efficient Retrieval Representation Fusion
by: Wu, Shangyu, et al.
Published: (2024)
by: Wu, Shangyu, et al.
Published: (2024)
KIF: Knowledge Identification and Fusion for Language Model Continual Learning
by: Feng, Yujie, et al.
Published: (2024)
by: Feng, Yujie, et al.
Published: (2024)
Similar Items
-
Safety Tax: Safety Alignment Makes Your Large Reasoning Models Less Reasonable
by: Huang, Tiansheng, et al.
Published: (2025) -
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
by: Huang, Tiansheng, et al.
Published: (2024) -
Virus: Harmful Fine-tuning Attack for Large Language Models Bypassing Guardrail Moderation
by: Huang, Tiansheng, et al.
Published: (2025) -
A Survey on Large Language Model-Based Game Agents
by: Hu, Sihao, et al.
Published: (2024) -
LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity
by: Tekin, Selim Furkan, et al.
Published: (2024)