Refining Positive and Toxic Samples for Dual Safety Self-Alignment of LLMs with Minimal Human Interventions
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Jingxin, Nan, Guoshun, Guan, Sheng, Leng, Sicong, Liu, Yilian, Wang, Zixiao, Ma, Yuyang, Zhou, Zhili, Hou, Yanzhao, Tao, Xiaofeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
por: Lu, Haolang, et al.
Publicado: (2025)
por: Lu, Haolang, et al.
Publicado: (2025)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
por: Lin, Honglin, et al.
Publicado: (2024)
por: Lin, Honglin, et al.
Publicado: (2024)
A Novel Indicator for Quantifying and Minimizing Information Utility Loss of Robot Teams
por: Zhao, Xiyu, et al.
Publicado: (2025)
por: Zhao, Xiyu, et al.
Publicado: (2025)
Adaptive Federated Learning in Heterogeneous Wireless Networks with Independent Sampling
por: Geng, Jiaxiang, et al.
Publicado: (2024)
por: Geng, Jiaxiang, et al.
Publicado: (2024)
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
por: Liu, Yilian, et al.
Publicado: (2026)
por: Liu, Yilian, et al.
Publicado: (2026)
Two Is Better Than One: Rotations Scale LoRAs
por: Guo, Hongcan, et al.
Publicado: (2025)
por: Guo, Hongcan, et al.
Publicado: (2025)
Adaptive Federated LoRA in Heterogeneous Wireless Networks with Independent Sampling
por: Hou, Yanzhao, et al.
Publicado: (2025)
por: Hou, Yanzhao, et al.
Publicado: (2025)
Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
por: Cao, Xinye, et al.
Publicado: (2025)
por: Cao, Xinye, et al.
Publicado: (2025)
Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
por: Cao, Xinye, et al.
Publicado: (2025)
por: Cao, Xinye, et al.
Publicado: (2025)
Advancing Expert Specialization for Better MoE
por: Guo, Hongcan, et al.
Publicado: (2025)
por: Guo, Hongcan, et al.
Publicado: (2025)
From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
por: Du, Hang, et al.
Publicado: (2025)
por: Du, Hang, et al.
Publicado: (2025)
Can We Improve Channel Reciprocity via Loop-back Compensation for RIS-assisted Physical Layer Key Generation
por: Xu, Ningya, et al.
Publicado: (2024)
por: Xu, Ningya, et al.
Publicado: (2024)
Minimizing Human Intervention in Online Classification
por: Réveillard, William, et al.
Publicado: (2025)
por: Réveillard, William, et al.
Publicado: (2025)
Stay Positive: Neural Refinement of Sample Weights
por: Nachman, Benjamin, et al.
Publicado: (2025)
por: Nachman, Benjamin, et al.
Publicado: (2025)
Benign Samples Matter! Fine-tuning On Outlier Benign Samples Severely Breaks Safety
por: Guan, Zihan, et al.
Publicado: (2025)
por: Guan, Zihan, et al.
Publicado: (2025)
Ocurrencia de granizos en Camagüey, su relación con la isoterma de 0 0 C del bulbo húmedo
por: Yilian Martínez Rodríguez
Publicado: (2011)
por: Yilian Martínez Rodríguez
Publicado: (2011)
EVALUACIÓN DE PROYECTOS DE I+D Y LA GESTIÓN DEL CONOCIMIENTO. IMPORTANCIA DE SU VINCULACIÓN PARA EL DESARROLLO ORGANIZACIONAL
por: Yilian Rodríguez-Clavijo
Publicado: (2012)
por: Yilian Rodríguez-Clavijo
Publicado: (2012)
Guía de evaluación estética de la sonrisa en ortodoncia
por: Yilian Pérez Mira
Publicado: (2022)
por: Yilian Pérez Mira
Publicado: (2022)
Estrategias de crecimiento empresarial aplicadas por hipermercados
por: Yilian Cefalá Chirinos
Publicado: (2003)
por: Yilian Cefalá Chirinos
Publicado: (2003)
ESTRATEGIA PARA LA GESTIÓN DE PROYECTOS DE COOPERACIÓN INTERNACIONAL EN UNA ENTIDAD DE CIENCIA E INNOVACIÓN TECNOLÓGICA
por: Yilian Rodríguez-Clavijo
Publicado: (2010)
por: Yilian Rodríguez-Clavijo
Publicado: (2010)
Condiciones termodinámicas asociadas a la ocurrencia de granizos en Camagüey
por: Yilian Martínez Rodríguez
Publicado: (2011)
por: Yilian Martínez Rodríguez
Publicado: (2011)
Sample-Efficient Alignment for LLMs
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
Unlocking Recursive Thinking of LLMs: Alignment via Refinement
por: Zhang, Haoke, et al.
Publicado: (2025)
por: Zhang, Haoke, et al.
Publicado: (2025)
CIDR: A Cooperative Integrated Dynamic Refining Method for Minimal Feature Removal Problem
por: Chen, Qian, et al.
Publicado: (2023)
por: Chen, Qian, et al.
Publicado: (2023)
Equipping Sketch Patches with Context-Aware Positional Encoding for Graphic Sketch Representation
por: Zang, Sicong, et al.
Publicado: (2024)
por: Zang, Sicong, et al.
Publicado: (2024)
PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference
por: Ji, Jiaming, et al.
Publicado: (2024)
por: Ji, Jiaming, et al.
Publicado: (2024)
D$^3$R-DETR: DETR with Dual-Domain Density Refinement for Tiny Object Detection in Aerial Images
por: Wen, Zixiao, et al.
Publicado: (2026)
por: Wen, Zixiao, et al.
Publicado: (2026)
Disentangling Deception and Hallucination Failures in LLMs
por: Lu, Haolang, et al.
Publicado: (2026)
por: Lu, Haolang, et al.
Publicado: (2026)
Refining Strokes by Learning Offset Attributes between Strokes for Flexible Sketch Edit at Stroke-Level
por: Zang, Sicong, et al.
Publicado: (2026)
por: Zang, Sicong, et al.
Publicado: (2026)
GRAFT: Grid-Aware Load Forecasting with Multi-Source Textual Alignment and Fusion
por: Lin, Fangzhou, et al.
Publicado: (2025)
por: Lin, Fangzhou, et al.
Publicado: (2025)
Inductive Convolution Nuclear Norm Minimization for Tensor Completion with Arbitrary Sampling
por: Li, Wei, et al.
Publicado: (2026)
por: Li, Wei, et al.
Publicado: (2026)
Valoración de las aguas de escurrimiento superficial, en ecosistemas forestales, en Pinar del Río
por: Yilian M. Morejón Miranda
Publicado: (2013)
por: Yilian M. Morejón Miranda
Publicado: (2013)
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
por: Zhang, Zhibo, et al.
Publicado: (2025)
por: Zhang, Zhibo, et al.
Publicado: (2025)
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
por: Singh, Himanshu, et al.
Publicado: (2026)
por: Singh, Himanshu, et al.
Publicado: (2026)
CARE: Decoding Time Safety Alignment via Rollback and Introspection Intervention
por: Hu, Xiaomeng, et al.
Publicado: (2025)
por: Hu, Xiaomeng, et al.
Publicado: (2025)
Differentiated Directional Intervention A Framework for Evading LLM Safety Alignment
por: Zhang, Peng, et al.
Publicado: (2025)
por: Zhang, Peng, et al.
Publicado: (2025)
Improving LLM Safety Alignment with Dual-Objective Optimization
por: Zhao, Xuandong, et al.
Publicado: (2025)
por: Zhao, Xuandong, et al.
Publicado: (2025)
Minimal Intervention Shared Control with Guaranteed Safety under Non-Convex Constraints
por: Chaubey, Shivam, et al.
Publicado: (2025)
por: Chaubey, Shivam, et al.
Publicado: (2025)
Positive Alignment: Artificial Intelligence for Human Flourishing
por: Laukkonen, Ruben, et al.
Publicado: (2026)
por: Laukkonen, Ruben, et al.
Publicado: (2026)
Position: Towards Bidirectional Human-AI Alignment
por: Shen, Hua, et al.
Publicado: (2024)
por: Shen, Hua, et al.
Publicado: (2024)
Ejemplares similares
-
Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
por: Lu, Haolang, et al.
Publicado: (2025) -
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
por: Lin, Honglin, et al.
Publicado: (2024) -
A Novel Indicator for Quantifying and Minimizing Information Utility Loss of Robot Teams
por: Zhao, Xiyu, et al.
Publicado: (2025) -
Adaptive Federated Learning in Heterogeneous Wireless Networks with Independent Sampling
por: Geng, Jiaxiang, et al.
Publicado: (2024) -
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
por: Liu, Yilian, et al.
Publicado: (2026)