Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning
Fuente:
arXiv
Guardado en:
| Autores principales: | Jan, Essa, AlDahoul, Nouar, Ali, Moiz, Ahmad, Faizan, Zaffar, Fareed, Zaki, Yasir |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
por: AlDahoul, Nouar, et al.
Publicado: (2025)
por: AlDahoul, Nouar, et al.
Publicado: (2025)
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
por: AlDahoul, Nouar, et al.
Publicado: (2025)
por: AlDahoul, Nouar, et al.
Publicado: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
por: AlDahoul, Nouar, et al.
Publicado: (2025)
por: AlDahoul, Nouar, et al.
Publicado: (2025)
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
por: Jan, Essa, et al.
Publicado: (2025)
por: Jan, Essa, et al.
Publicado: (2025)
A Longitudinal Analysis of Racial and Gender Bias in New York Times and Fox News Images and Articles
por: Ibrahim, Hazem, et al.
Publicado: (2024)
por: Ibrahim, Hazem, et al.
Publicado: (2024)
AI-generated faces influence gender stereotypes and racial homogenization
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Self-Reflection Makes Large Language Models Safer, Less Biased, and Ideologically Neutral
por: Liu, Fengyuan, et al.
Publicado: (2024)
por: Liu, Fengyuan, et al.
Publicado: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Neutralizing the Narrative: AI-Powered Debiasing of Online News Articles
por: Kuo, Chen Wei, et al.
Publicado: (2025)
por: Kuo, Chen Wei, et al.
Publicado: (2025)
Fine-tuned Vision Language Model for Localization of Parasitic Eggs in Microscopic Images
por: Sien, Chan Hao, et al.
Publicado: (2026)
por: Sien, Chan Hao, et al.
Publicado: (2026)
Inclusive content reduces racial and gender biases, yet non-inclusive content dominates popular culture
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Exploring Vision Language Models for Facial Attribute Recognition: Emotion, Race, Gender, and Age
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Enhancing Password Security Through a High-Accuracy Scoring Framework Using Random Forests
por: Mazelan, Muhammed El Mustaqeem, et al.
Publicado: (2025)
por: Mazelan, Muhammed El Mustaqeem, et al.
Publicado: (2025)
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma
por: AlDahoul, Nouar, et al.
Publicado: (2024)
por: AlDahoul, Nouar, et al.
Publicado: (2024)
Real-Time Human Detection for Aerial Captured Video Sequences via Deep Models
por: AlDahoul, Nouar, et al.
Publicado: (2026)
por: AlDahoul, Nouar, et al.
Publicado: (2026)
Schadenfreude in the Digital Public Sphere: A cross-national and decade-long analysis of Facebook news engagement
por: Aldahoul, Nouar, et al.
Publicado: (2026)
por: Aldahoul, Nouar, et al.
Publicado: (2026)
Semantic-Aware Advanced Persistent Threat Detection Using Autoencoders on LLM-Encoded System Logs
por: Mohammed, Waleed Khan, et al.
Publicado: (2026)
por: Mohammed, Waleed Khan, et al.
Publicado: (2026)
Fine-tuning and Utilization Methods of Domain-specific LLMs
por: Jeong, Cheonsu
Publicado: (2024)
por: Jeong, Cheonsu
Publicado: (2024)
An explainable Recursive Feature Elimination to detect Advanced Persistent Threats using Random Forest classifier
por: Mutalib, Noor Hazlina Abdul, et al.
Publicado: (2025)
por: Mutalib, Noor Hazlina Abdul, et al.
Publicado: (2025)
A Conceptual Exploration of Generative AI-Induced Cognitive Dissonance and its Emergence in University-Level Academic Writing
por: Seran, Carl Errol, et al.
Publicado: (2025)
por: Seran, Carl Errol, et al.
Publicado: (2025)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
por: Wang, Yibo, et al.
Publicado: (2025)
por: Wang, Yibo, et al.
Publicado: (2025)
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
por: Faisal, Faizan
Publicado: (2026)
por: Faisal, Faizan
Publicado: (2026)
Aloe: A Family of Fine-tuned Open Healthcare LLMs
por: Gururajan, Ashwin Kumar, et al.
Publicado: (2024)
por: Gururajan, Ashwin Kumar, et al.
Publicado: (2024)
Large Language Models are often politically extreme, usually ideologically inconsistent, and persuasive even in informational contexts
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
ITERTL: An Iterative Framework for Fine-tuning LLMs for RTL Code Generation
por: Wu, Peiyang, et al.
Publicado: (2024)
por: Wu, Peiyang, et al.
Publicado: (2024)
Mitigating Clickbait: An Approach to Spoiler Generation Using Multitask Learning
por: Pal, Sayantan, et al.
Publicado: (2024)
por: Pal, Sayantan, et al.
Publicado: (2024)
Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
por: Kim, Kyeonghyun, et al.
Publicado: (2025)
por: Kim, Kyeonghyun, et al.
Publicado: (2025)
Phonetic Perturbations Reveal Tokenizer-Rooted Safety Gaps in LLMs
por: Aswal, Darpan, et al.
Publicado: (2025)
por: Aswal, Darpan, et al.
Publicado: (2025)
MuLD: The Multitask Long Document Benchmark
por: Hudson, G Thomas, et al.
Publicado: (2022)
por: Hudson, G Thomas, et al.
Publicado: (2022)
Safety Anchor: Defending Harmful Fine-tuning via Geometric Bottlenecks
por: Lu, Guoxin, et al.
Publicado: (2026)
por: Lu, Guoxin, et al.
Publicado: (2026)
Affordably Fine-tuned LLMs Provide Better Answers to Course-specific MCQs
por: Raimondi, Bianca, et al.
Publicado: (2025)
por: Raimondi, Bianca, et al.
Publicado: (2025)
HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
por: Vasilatos, Christoforos, et al.
Publicado: (2023)
por: Vasilatos, Christoforos, et al.
Publicado: (2023)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
por: Mao, Yingzhi, et al.
Publicado: (2025)
por: Mao, Yingzhi, et al.
Publicado: (2025)
Fine-tuned Large Language Models (LLMs): Improved Prompt Injection Attacks Detection
por: Rahman, Md Abdur, et al.
Publicado: (2024)
por: Rahman, Md Abdur, et al.
Publicado: (2024)
Identity Lock: Locking API Fine-tuned LLMs With Identity-based Wake Words
por: Su, Hongyu, et al.
Publicado: (2025)
por: Su, Hongyu, et al.
Publicado: (2025)
Fundamental Safety-Capability Trade-offs in Fine-tuning Large Language Models
por: Chen, Pin-Yu, et al.
Publicado: (2025)
por: Chen, Pin-Yu, et al.
Publicado: (2025)
RLSF: Fine-tuning LLMs via Symbolic Feedback
por: Jha, Piyush, et al.
Publicado: (2024)
por: Jha, Piyush, et al.
Publicado: (2024)
LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
por: Ren, Qibing, et al.
Publicado: (2024)
por: Ren, Qibing, et al.
Publicado: (2024)
Ejemplares similares
-
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
por: AlDahoul, Nouar, et al.
Publicado: (2025) -
Benchmarking the Medical Understanding and Reasoning of Large Language Models in Arabic Healthcare Tasks
por: AlDahoul, Nouar, et al.
Publicado: (2025) -
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
por: AlDahoul, Nouar, et al.
Publicado: (2025) -
A Novel BERT-based Classifier to Detect Political Leaning of YouTube Videos based on their Titles
por: AlDahoul, Nouar, et al.
Publicado: (2024) -
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
por: Jan, Essa, et al.
Publicado: (2025)