Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Taraghi, Mina, Pequignot, Yann, Nikanjam, Amin, Merzouk, Mohamed Amine, Khomh, Foutse |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
por: Majdinasab, Vahid, et al.
Publicado: (2025)
por: Majdinasab, Vahid, et al.
Publicado: (2025)
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
por: Majdinasab, Vahid, et al.
Publicado: (2024)
por: Majdinasab, Vahid, et al.
Publicado: (2024)
ReCatcher: Towards LLMs Regression Testing for Code Generation
por: Abbassi, Altaf Allah, et al.
Publicado: (2025)
por: Abbassi, Altaf Allah, et al.
Publicado: (2025)
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
por: Morovati, Mohammad Mehdi, et al.
Publicado: (2023)
por: Morovati, Mohammad Mehdi, et al.
Publicado: (2023)
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
por: Tambon, Florian, et al.
Publicado: (2024)
por: Tambon, Florian, et al.
Publicado: (2024)
Adversarial Moral Stress Testing of Large Language Models
por: Jamshidi, Saeid, et al.
Publicado: (2026)
por: Jamshidi, Saeid, et al.
Publicado: (2026)
Bugs in Large Language Models Generated Code: An Empirical Study
por: Tambon, Florian, et al.
Publicado: (2024)
por: Tambon, Florian, et al.
Publicado: (2024)
Real Faults in Model Context Protocol (MCP) Software: a Comprehensive Taxonomy
por: Taraghi, Mina, et al.
Publicado: (2026)
por: Taraghi, Mina, et al.
Publicado: (2026)
Secure Tool Manifest and Digital Signing Solution for Verifiable MCP and LLM Pipelines
por: Jamshidi, Saeid, et al.
Publicado: (2026)
por: Jamshidi, Saeid, et al.
Publicado: (2026)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
por: Majdinasab, Vahid, et al.
Publicado: (2024)
por: Majdinasab, Vahid, et al.
Publicado: (2024)
Toward Debugging Deep Reinforcement Learning Programs with RLExplorer
por: Bouchoucha, Rached, et al.
Publicado: (2024)
por: Bouchoucha, Rached, et al.
Publicado: (2024)
An Efficient Model Maintenance Approach for MLOps
por: Majidi, Forough, et al.
Publicado: (2024)
por: Majidi, Forough, et al.
Publicado: (2024)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
por: Da Silva, Leuson, et al.
Publicado: (2024)
por: Da Silva, Leuson, et al.
Publicado: (2024)
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
por: Le, Linh, et al.
Publicado: (2026)
por: Le, Linh, et al.
Publicado: (2026)
Fault Localization in Deep Learning-based Software: A System-level Approach
por: Morovati, Mohammad Mehdi, et al.
Publicado: (2024)
por: Morovati, Mohammad Mehdi, et al.
Publicado: (2024)
Machine Learning Robustness: A Primer
por: Braiek, Houssem Ben, et al.
Publicado: (2024)
por: Braiek, Houssem Ben, et al.
Publicado: (2024)
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4
por: Abukhalaf, Seif, et al.
Publicado: (2024)
por: Abukhalaf, Seif, et al.
Publicado: (2024)
RefAgent: A Multi-agent LLM-based Framework for Automatic Software Refactoring
por: Oueslati, Khouloud, et al.
Publicado: (2025)
por: Oueslati, Khouloud, et al.
Publicado: (2025)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
por: Ngassom, Sylvain Kouemo, et al.
Publicado: (2024)
por: Ngassom, Sylvain Kouemo, et al.
Publicado: (2024)
Deep Learning Model Reuse in the HuggingFace Community: Challenges, Benefit and Trends
por: Taraghi, Mina, et al.
Publicado: (2024)
por: Taraghi, Mina, et al.
Publicado: (2024)
Semantic Attacks on Tool-Augmented LLMs: Securing the Model Context Protocol Against Descriptor-Level Manipulation
por: Jamshidi, Saeid, et al.
Publicado: (2025)
por: Jamshidi, Saeid, et al.
Publicado: (2025)
What Information Contributes to Log-based Anomaly Detection? Insights from a Configurable Transformer-Based Approach
por: Wu, Xingfang, et al.
Publicado: (2024)
por: Wu, Xingfang, et al.
Publicado: (2024)
From LLMs to Edge: Parameter-Efficient Fine-Tuning on Edge Devices
por: Slamanig, Georg, et al.
Publicado: (2025)
por: Slamanig, Georg, et al.
Publicado: (2025)
Leveraging Machine Learning Techniques in Intrusion Detection Systems for Internet of Things
por: Jamshidi, Saeid, et al.
Publicado: (2025)
por: Jamshidi, Saeid, et al.
Publicado: (2025)
Evaluating Machine Learning-Driven Intrusion Detection Systems in IoT: Performance and Energy Consumption
por: Jamshidi, Saeid, et al.
Publicado: (2025)
por: Jamshidi, Saeid, et al.
Publicado: (2025)
SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs
por: Aghili, Roozbeh, et al.
Publicado: (2025)
por: Aghili, Roozbeh, et al.
Publicado: (2025)
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
por: Bhosale, Mahesh, et al.
Publicado: (2026)
por: Bhosale, Mahesh, et al.
Publicado: (2026)
Understanding and Preserving Safety in Fine-Tuned LLMs
por: Zhang, Jiawen, et al.
Publicado: (2026)
por: Zhang, Jiawen, et al.
Publicado: (2026)
Defending MoE LLMs against Harmful Fine-Tuning via Safety Routing Alignment
por: Kim, Jaehan, et al.
Publicado: (2025)
por: Kim, Jaehan, et al.
Publicado: (2025)
Parameter-Efficient Fine-Tuning With Adapters
por: Chen, Keyu, et al.
Publicado: (2024)
por: Chen, Keyu, et al.
Publicado: (2024)
Analysing Safety Risks in LLMs Fine-Tuned with Pseudo-Malicious Cyber Security Data
por: ElZemity, Adel, et al.
Publicado: (2025)
por: ElZemity, Adel, et al.
Publicado: (2025)
NaRA: Noise-Aware LoRA for Parameter-Efficient Fine-Tuning of Diffusion LLMs
por: Wang, Shuaidi, et al.
Publicado: (2026)
por: Wang, Shuaidi, et al.
Publicado: (2026)
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
por: Giordani, Jeremiah
Publicado: (2025)
por: Giordani, Jeremiah
Publicado: (2025)
NeuronTune: Fine-Grained Neuron Modulation for Balanced Safety-Utility Alignment in LLMs
por: Pan, Birong, et al.
Publicado: (2025)
por: Pan, Birong, et al.
Publicado: (2025)
Partial Order in Chaos: Consensus on Feature Attributions in the Rashomon Set
por: Laberge, Gabriel, et al.
Publicado: (2021)
por: Laberge, Gabriel, et al.
Publicado: (2021)
Diffusion-based Adversarial Purification for Intrusion Detection
por: Merzouk, Mohamed Amine, et al.
Publicado: (2024)
por: Merzouk, Mohamed Amine, et al.
Publicado: (2024)
Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
por: Gulati, Idhant, et al.
Publicado: (2026)
por: Gulati, Idhant, et al.
Publicado: (2026)
Parameter Efficiency Is Not Memory Efficiency: Rethinking Fine-Tuning for On-Device LLM Adaptation
por: Tenison, Irene, et al.
Publicado: (2026)
por: Tenison, Irene, et al.
Publicado: (2026)
Securing Time in Energy IoT: A Clock-Dynamics-Aware Spatio-Temporal Graph Attention Network for Clock Drift Attacks and Y2K38 Failures
por: Jamshidi, Saeid, et al.
Publicado: (2026)
por: Jamshidi, Saeid, et al.
Publicado: (2026)
Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
por: Shah, Mehil B, et al.
Publicado: (2025)
por: Shah, Mehil B, et al.
Publicado: (2025)
Ejemplares similares
-
Prism: Dynamic and Flexible Benchmarking of LLMs Code Generation with Monte Carlo Tree Search
por: Majdinasab, Vahid, et al.
Publicado: (2025) -
DeepCodeProbe: Towards Understanding What Models Trained on Code Learn
por: Majdinasab, Vahid, et al.
Publicado: (2024) -
ReCatcher: Towards LLMs Regression Testing for Code Generation
por: Abbassi, Altaf Allah, et al.
Publicado: (2025) -
Common Challenges of Deep Reinforcement Learning Applications Development: An Empirical Study
por: Morovati, Mohammad Mehdi, et al.
Publicado: (2023) -
TaskEval: Assessing Difficulty of Code Generation Tasks for Large Language Models
por: Tambon, Florian, et al.
Publicado: (2024)