Scam Shield: Multi-Model Voting and Fine-Tuned LLMs Against Adversarial Attacks
Fuente:
arXiv
Guardado en:
| Autores principales: | Chang, Chen-Wei, Sarkar, Shailik, Salemi, Hossein, Kim, Hyungmin, Mitra, Shutonu, Purohit, Hemant, Zhang, Fengxiu, Hong, Michin, Cho, Jin-Hee, Lu, Chang-Tien |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
por: Chang, Chen-Wei, et al.
Publicado: (2024)
por: Chang, Chen-Wei, et al.
Publicado: (2024)
SCVI: Bridging Social and Cyber Dimensions for Comprehensive Vulnerability Assessment
por: Mitra, Shutonu, et al.
Publicado: (2025)
por: Mitra, Shutonu, et al.
Publicado: (2025)
Debiasing Large Language Models toward Social Factors in Online Behavior Analytics through Prompt Knowledge Tuning
por: Salemi, Hossein, et al.
Publicado: (2026)
por: Salemi, Hossein, et al.
Publicado: (2026)
MVeLMA: Multimodal Vegetation Loss Modeling Architecture for Predicting Post-fire Vegetation Loss
por: Ravi, Meenu, et al.
Publicado: (2025)
por: Ravi, Meenu, et al.
Publicado: (2025)
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)
por: Liu, Fan, et al.
Publicado: (2024)
ScamPilot: Simulating Conversations with LLMs to Protect Against Online Scams
por: Hoffman, Owen, et al.
Publicado: (2026)
por: Hoffman, Owen, et al.
Publicado: (2026)
Closing the Knowledge Gap in Designing Data Annotation Interfaces for AI-powered Disaster Management Analytic Systems
por: Ara, Zinat, et al.
Publicado: (2024)
por: Ara, Zinat, et al.
Publicado: (2024)
ShieldUp!: Inoculating Users Against Online Scams Using A Game Based Intervention
por: Roy, Abhishek, et al.
Publicado: (2025)
por: Roy, Abhishek, et al.
Publicado: (2025)
Adversarial Attacks Against Deep Learning-Based Radio Frequency Fingerprint Identification
por: Ma, Jie, et al.
Publicado: (2025)
por: Ma, Jie, et al.
Publicado: (2025)
Knowledge-guided Continual Learning for Behavioral Analytics Systems
por: Senarath, Yasas, et al.
Publicado: (2025)
por: Senarath, Yasas, et al.
Publicado: (2025)
Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks
por: Alexiou, Michail S., et al.
Publicado: (2026)
por: Alexiou, Michail S., et al.
Publicado: (2026)
Library Social Services: How Relevant Are Social Services to the Existing Library Missions?
por: Ayoung Yoon, et al.
Publicado: (2024)
por: Ayoung Yoon, et al.
Publicado: (2024)
Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs
por: Mitra, Subhadip
Publicado: (2026)
por: Mitra, Subhadip
Publicado: (2026)
Comparing Retrieval-Augmentation and Parameter-Efficient Fine-Tuning for Privacy-Preserving Personalization of Large Language Models
por: Salemi, Alireza, et al.
Publicado: (2024)
por: Salemi, Alireza, et al.
Publicado: (2024)
VEXA: Evidence-Grounded and Persona-Adaptive Explanations for Scam Risk Sensemaking
por: An, Heajun, et al.
Publicado: (2026)
por: An, Heajun, et al.
Publicado: (2026)
RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System
por: Choudhury, Nitin, et al.
Publicado: (2026)
por: Choudhury, Nitin, et al.
Publicado: (2026)
DeepVoting: Learning and Fine-Tuning Voting Rules with Canonical Embeddings
por: Matone, Leonardo, et al.
Publicado: (2024)
por: Matone, Leonardo, et al.
Publicado: (2024)
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks
por: Li, Rongchang, et al.
Publicado: (2024)
por: Li, Rongchang, et al.
Publicado: (2024)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
por: Feng, Weitao, et al.
Publicado: (2025)
por: Feng, Weitao, et al.
Publicado: (2025)
Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams
por: Basta, Nardine, et al.
Publicado: (2025)
por: Basta, Nardine, et al.
Publicado: (2025)
Self-Evaluation as a Defense Against Adversarial Attacks on LLMs
por: Brown, Hannah, et al.
Publicado: (2024)
por: Brown, Hannah, et al.
Publicado: (2024)
Resilient Multi-Agent Negotiation for Medical Supply Chains:Integrating LLMs and Blockchain for Transparent Coordination
por: ALMutairi, Mariam, et al.
Publicado: (2025)
por: ALMutairi, Mariam, et al.
Publicado: (2025)
Sharpening the Spear: Adaptive Expert-Guided Adversarial Attack Against DRL-based Autonomous Driving Policies
por: Fan, Junchao, et al.
Publicado: (2025)
por: Fan, Junchao, et al.
Publicado: (2025)
Who Can See Through You? Adversarial Shielding Against VLM-Based Attribute Inference Attacks
por: Fan, Yucheng, et al.
Publicado: (2025)
por: Fan, Yucheng, et al.
Publicado: (2025)
Enhancing Temporal Link Prediction with HierTKG: A Hierarchical Temporal Knowledge Graph Framework
por: Almutairi, Mariam, et al.
Publicado: (2024)
por: Almutairi, Mariam, et al.
Publicado: (2024)
PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
por: Sun, Zhen, et al.
Publicado: (2024)
por: Sun, Zhen, et al.
Publicado: (2024)
ORIS: Online Active Learning Using Reinforcement Learning-based Inclusive Sampling for Robust Streaming Analytics System
por: Pandey, Rahul, et al.
Publicado: (2024)
por: Pandey, Rahul, et al.
Publicado: (2024)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
por: Lee, Changhun, et al.
Publicado: (2024)
por: Lee, Changhun, et al.
Publicado: (2024)
ROAST: Risk-aware Outlier-exposure for Adversarial Selective Training of Anomaly Detectors Against Evasion Attacks
por: Elnawawy, Mohammed, et al.
Publicado: (2026)
por: Elnawawy, Mohammed, et al.
Publicado: (2026)
Persistent Backdoor Attacks under Continual Fine-Tuning of LLMs
por: Cui, Jing, et al.
Publicado: (2025)
por: Cui, Jing, et al.
Publicado: (2025)
Tuning for Two Adversaries: Enhancing the Robustness Against Transfer and Query-Based Attacks using Hyperparameter Tuning
por: Zimmer, Pascal, et al.
Publicado: (2025)
por: Zimmer, Pascal, et al.
Publicado: (2025)
Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
por: Chen, Zhiyang, et al.
Publicado: (2025)
por: Chen, Zhiyang, et al.
Publicado: (2025)
Defending Against Weight-Poisoning Backdoor Attacks for Parameter-Efficient Fine-Tuning
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
GreedyPixel: Fine-Grained Black-Box Adversarial Attack Via Greedy Algorithm
por: Wang, Hanrui, et al.
Publicado: (2025)
por: Wang, Hanrui, et al.
Publicado: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
por: Poppi, Samuele, et al.
Publicado: (2024)
por: Poppi, Samuele, et al.
Publicado: (2024)
AI-KD: Adversarial learning and Implicit regularization for self-Knowledge Distillation
por: Kim, Hyungmin, et al.
Publicado: (2022)
por: Kim, Hyungmin, et al.
Publicado: (2022)
Can Differentially Private Fine-tuning LLMs Protect Against Privacy Attacks?
por: Du, Hao, et al.
Publicado: (2025)
por: Du, Hao, et al.
Publicado: (2025)
SCOUT: A Defense Against Data Poisoning Attacks in Fine-Tuned Language Models
por: Afane, Mohamed, et al.
Publicado: (2025)
por: Afane, Mohamed, et al.
Publicado: (2025)
Membership Inference Attacks for Face Images Against Fine-Tuned Latent Diffusion Models
por: Holme, Lauritz Christian, et al.
Publicado: (2025)
por: Holme, Lauritz Christian, et al.
Publicado: (2025)
Patrol Security Game: Defending Against Adversary with Freedom in Attack Timing, Location, and Duration
por: Yang, Hao-Tsung, et al.
Publicado: (2024)
por: Yang, Hao-Tsung, et al.
Publicado: (2024)
Ejemplares similares
-
Exposing LLM Vulnerabilities: Adversarial Scam Detection and Performance
por: Chang, Chen-Wei, et al.
Publicado: (2024) -
SCVI: Bridging Social and Cyber Dimensions for Comprehensive Vulnerability Assessment
por: Mitra, Shutonu, et al.
Publicado: (2025) -
Debiasing Large Language Models toward Social Factors in Online Behavior Analytics through Prompt Knowledge Tuning
por: Salemi, Hossein, et al.
Publicado: (2026) -
MVeLMA: Multimodal Vegetation Loss Modeling Architecture for Predicting Post-fire Vegetation Loss
por: Ravi, Meenu, et al.
Publicado: (2025) -
Adversarial Tuning: Defending Against Jailbreak Attacks for LLMs
por: Liu, Fan, et al.
Publicado: (2024)