FFT: Towards Harmlessness Evaluation and Analysis for LLMs with Factuality, Fairness, Toxicity
Fuente:
arXiv
Guardado en:
| Autores principales: | Cui, Shiyao, Zhang, Zhenyu, Chen, Yilong, Zhang, Wenyuan, Liu, Tianyun, Wang, Siqi, Liu, Tingwen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
por: Xie, Yuanbo, et al.
Publicado: (2025)
por: Xie, Yuanbo, et al.
Publicado: (2025)
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
por: Li, Lijun, et al.
Publicado: (2025)
por: Li, Lijun, et al.
Publicado: (2025)
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
por: Hastuti, Rochana Prih, et al.
Publicado: (2025)
por: Hastuti, Rochana Prih, et al.
Publicado: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
por: Xing, Wenpeng, et al.
Publicado: (2025)
por: Xing, Wenpeng, et al.
Publicado: (2025)
Safety Alignment Should Be Made More Than Just A Few Attention Heads
por: Huang, Chao, et al.
Publicado: (2025)
por: Huang, Chao, et al.
Publicado: (2025)
Jailbreaking LLMs via Semantically Relevant Nested Scenarios with Targeted Toxic Knowledge
por: Xu, Ning, et al.
Publicado: (2025)
por: Xu, Ning, et al.
Publicado: (2025)
Detecting RAG Extraction Attack via Dual-Path Runtime Integrity Game
por: Xie, Yuanbo, et al.
Publicado: (2026)
por: Xie, Yuanbo, et al.
Publicado: (2026)
S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models
por: Yuan, Xiaohan, et al.
Publicado: (2024)
por: Yuan, Xiaohan, et al.
Publicado: (2024)
Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMs
por: Fastowski, Alina, et al.
Publicado: (2025)
por: Fastowski, Alina, et al.
Publicado: (2025)
How Real is Your Jailbreak? Fine-grained Jailbreak Evaluation with Anchored Reference
por: Liu, Songyang, et al.
Publicado: (2026)
por: Liu, Songyang, et al.
Publicado: (2026)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
por: Liu, Songyang, et al.
Publicado: (2025)
por: Liu, Songyang, et al.
Publicado: (2025)
Toward Copyright Integrity and Verifiability via Multi-Bit Watermarking for Intelligent Transportation Systems
por: Wang, Yihao, et al.
Publicado: (2025)
por: Wang, Yihao, et al.
Publicado: (2025)
PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
por: Huo, Jiahao, et al.
Publicado: (2025)
por: Huo, Jiahao, et al.
Publicado: (2025)
Towards Understanding the Safety Boundaries of DeepSeek Models: Evaluation and Findings
por: Ying, Zonghao, et al.
Publicado: (2025)
por: Ying, Zonghao, et al.
Publicado: (2025)
From Compression to Accountability: Harmless Copyright Protection for Dataset Distillation
por: Liang, Yan, et al.
Publicado: (2026)
por: Liang, Yan, et al.
Publicado: (2026)
Jailbreaking Commercial Black-Box LLMs with Explicitly Harmful Prompts
por: Zhang, Chiyu, et al.
Publicado: (2025)
por: Zhang, Chiyu, et al.
Publicado: (2025)
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
por: Zhang, Zhexin, et al.
Publicado: (2024)
por: Zhang, Zhexin, et al.
Publicado: (2024)
Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models
por: Sakib, Shahnewaz Karim, et al.
Publicado: (2025)
por: Sakib, Shahnewaz Karim, et al.
Publicado: (2025)
The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training
por: Zhang, Rui, et al.
Publicado: (2026)
por: Zhang, Rui, et al.
Publicado: (2026)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
por: Gu, Tianle, et al.
Publicado: (2025)
por: Gu, Tianle, et al.
Publicado: (2025)
PREE: Towards Harmless and Adaptive Fingerprint Editing in Large Language Models via Knowledge Prefix Enhancement
por: Yue, Xubin, et al.
Publicado: (2025)
por: Yue, Xubin, et al.
Publicado: (2025)
Self and Cross-Model Distillation for LLMs: Effective Methods for Refusal Pattern Alignment
por: Li, Jie, et al.
Publicado: (2024)
por: Li, Jie, et al.
Publicado: (2024)
Towards Safe AI Clinicians: A Comprehensive Study on Large Language Model Jailbreaking in Healthcare
por: Zhang, Hang, et al.
Publicado: (2025)
por: Zhang, Hang, et al.
Publicado: (2025)
Overlooked Safety Vulnerability in LLMs: Malicious Intelligent Optimization Algorithm Request and its Jailbreak
por: Gu, Haoran, et al.
Publicado: (2026)
por: Gu, Haoran, et al.
Publicado: (2026)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
por: Liu, Yepeng, et al.
Publicado: (2025)
por: Liu, Yepeng, et al.
Publicado: (2025)
TRUCE: Private Benchmarking to Prevent Contamination and Improve Comparative Evaluation of LLMs
por: Rajore, Tanmay, et al.
Publicado: (2024)
por: Rajore, Tanmay, et al.
Publicado: (2024)
Semantic-Preserving Adversarial Attacks on LLMs: An Adaptive Greedy Binary Search Approach
por: Zhang, Chong, et al.
Publicado: (2025)
por: Zhang, Chong, et al.
Publicado: (2025)
The Model's Language Matters: A Comparative Privacy Analysis of LLMs
por: Mishra, Abhishek K., et al.
Publicado: (2025)
por: Mishra, Abhishek K., et al.
Publicado: (2025)
Lightweight Yet Secure: Secure Scripting Language Generation via Lightweight LLMs
por: Zhang, Keyang, et al.
Publicado: (2026)
por: Zhang, Keyang, et al.
Publicado: (2026)
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
por: Yan, Yu, et al.
Publicado: (2025)
por: Yan, Yu, et al.
Publicado: (2025)
Efficient and Stealthy Jailbreak Attacks via Adversarial Prompt Distillation from LLMs to SLMs
por: Li, Xiang, et al.
Publicado: (2025)
por: Li, Xiang, et al.
Publicado: (2025)
Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs
por: Chen, Yunhao, et al.
Publicado: (2025)
por: Chen, Yunhao, et al.
Publicado: (2025)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
por: He, Yu, et al.
Publicado: (2025)
por: He, Yu, et al.
Publicado: (2025)
One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs
por: Li, Linbao, et al.
Publicado: (2025)
por: Li, Linbao, et al.
Publicado: (2025)
More Haste, Less Speed: Weaker Single-Layer Watermark Improves Distortion-Free Watermark Ensembles
por: Chen, Ruibo, et al.
Publicado: (2026)
por: Chen, Ruibo, et al.
Publicado: (2026)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
por: Zhu, Kaijie, et al.
Publicado: (2023)
por: Zhu, Kaijie, et al.
Publicado: (2023)
ShadowCoT: Cognitive Hijacking for Stealthy Reasoning Backdoors in LLMs
por: Zhao, Gejian, et al.
Publicado: (2025)
por: Zhao, Gejian, et al.
Publicado: (2025)
MGTEVAL: An Interactive Platform for Systemtic Evaluation of Machine-Generated Text Detectors
por: Li, Yuanfan, et al.
Publicado: (2026)
por: Li, Yuanfan, et al.
Publicado: (2026)
Harmless Backdoor-based Client-side Watermarking in Federated Learning
por: Luo, Kaijing, et al.
Publicado: (2024)
por: Luo, Kaijing, et al.
Publicado: (2024)
Stealthy Backdoor Attacks against LLMs Based on Natural Style Triggers
por: Wei, Jiali, et al.
Publicado: (2026)
por: Wei, Jiali, et al.
Publicado: (2026)
Ejemplares similares
-
Beyond Surface Alignment: Rebuilding LLMs Safety Mechanism via Probabilistically Ablating Refusal Direction
por: Xie, Yuanbo, et al.
Publicado: (2025) -
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation
por: Li, Lijun, et al.
Publicado: (2025) -
Factuality Beyond Coherence: Evaluating LLM Watermarking Methods for Medical Texts
por: Hastuti, Rochana Prih, et al.
Publicado: (2025) -
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
por: Xing, Wenpeng, et al.
Publicado: (2025) -
Safety Alignment Should Be Made More Than Just A Few Attention Heads
por: Huang, Chao, et al.
Publicado: (2025)