DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Ou, Haoran, Chen, Kangjie, Deng, Gelei, Liu, Hangcheng, Zhang, Jie, Zhang, Tianwei, Lam, Kwok-Yan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
di: Ou, Haoran, et al.
Pubblicazione: (2025)
di: Ou, Haoran, et al.
Pubblicazione: (2025)
Oedipus: LLM-enchanced Reasoning CAPTCHA Solver
di: Deng, Gelei, et al.
Pubblicazione: (2024)
di: Deng, Gelei, et al.
Pubblicazione: (2024)
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
di: Zhao, Shiqian, et al.
Pubblicazione: (2025)
di: Zhao, Shiqian, et al.
Pubblicazione: (2025)
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
di: Jiang, Hanrui, et al.
Pubblicazione: (2025)
di: Jiang, Hanrui, et al.
Pubblicazione: (2025)
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
di: Yang, Xianglin, et al.
Pubblicazione: (2026)
di: Yang, Xianglin, et al.
Pubblicazione: (2026)
BadLingual: A Novel Lingual-Backdoor Attack against Large Language Models
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
IRCopilot: Automated Incident Response with Large Language Models
di: Lin, Xihuan, et al.
Pubblicazione: (2025)
di: Lin, Xihuan, et al.
Pubblicazione: (2025)
Adversarial Attacks Against Automated Fact-Checking: A Survey
di: Liu, Fanzhen, et al.
Pubblicazione: (2025)
di: Liu, Fanzhen, et al.
Pubblicazione: (2025)
What Makes a Good LLM Agent for Real-world Penetration Testing?
di: Deng, Gelei, et al.
Pubblicazione: (2026)
di: Deng, Gelei, et al.
Pubblicazione: (2026)
LLM-Driven Feature-Level Adversarial Attacks on Android Malware Detectors
di: Lan, Tianwei, et al.
Pubblicazione: (2025)
di: Lan, Tianwei, et al.
Pubblicazione: (2025)
ART: Automatic Red-teaming for Text-to-Image Models to Protect Benign Users
di: Li, Guanlin, et al.
Pubblicazione: (2024)
di: Li, Guanlin, et al.
Pubblicazione: (2024)
Pandora: Jailbreak GPTs by Retrieval Augmented Generation Poisoning
di: Deng, Gelei, et al.
Pubblicazione: (2024)
di: Deng, Gelei, et al.
Pubblicazione: (2024)
SafeRedir: Prompt Embedding Redirection for Robust Unlearning in Image Generation Models
di: Liu, Renyang, et al.
Pubblicazione: (2026)
di: Liu, Renyang, et al.
Pubblicazione: (2026)
Mind Your HEARTBEAT! Claw Background Execution Inherently Enables Silent Memory Pollution
di: Zhang, Yechao, et al.
Pubblicazione: (2026)
di: Zhang, Yechao, et al.
Pubblicazione: (2026)
CipherGuard: Compiler-aided Mitigation against Ciphertext Side-channel Attacks
di: Jiang, Ke, et al.
Pubblicazione: (2025)
di: Jiang, Ke, et al.
Pubblicazione: (2025)
PentestEval: Benchmarking LLM-based Penetration Testing with Modular and Stage-Level Design
di: Yang, Ruozhao, et al.
Pubblicazione: (2025)
di: Yang, Ruozhao, et al.
Pubblicazione: (2025)
Enhancing Model Defense Against Jailbreaks with Proactive Safety Reasoning
di: Yang, Xianglin, et al.
Pubblicazione: (2025)
di: Yang, Xianglin, et al.
Pubblicazione: (2025)
ObfusBFA: A Holistic Approach to Safeguarding DNNs from Different Types of Bit-Flip Attacks
di: Yan, Xiaobei, et al.
Pubblicazione: (2025)
di: Yan, Xiaobei, et al.
Pubblicazione: (2025)
Robust-Wide: Robust Watermarking against Instruction-driven Image Editing
di: Hu, Runyi, et al.
Pubblicazione: (2024)
di: Hu, Runyi, et al.
Pubblicazione: (2024)
BitHydra: Towards Bit-flip Inference Cost Attack against Large Language Models
di: Yan, Xiaobei, et al.
Pubblicazione: (2025)
di: Yan, Xiaobei, et al.
Pubblicazione: (2025)
Prompt Injection attack against LLM-integrated Applications
di: Liu, Yi, et al.
Pubblicazione: (2023)
di: Liu, Yi, et al.
Pubblicazione: (2023)
Jailbreaking Prompt Attack: A Controllable Adversarial Attack against Diffusion Models
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
di: Ma, Jiachen, et al.
Pubblicazione: (2024)
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
di: Deng, Gelei, et al.
Pubblicazione: (2023)
di: Deng, Gelei, et al.
Pubblicazione: (2023)
Membership Inference Attacks Against Video Large Language Models
di: Song, Wei, et al.
Pubblicazione: (2026)
di: Song, Wei, et al.
Pubblicazione: (2026)
The Gradient Puppeteer: Adversarial Domination in Gradient Leakage Attacks through Model Poisoning
di: Xiang, Kunlan, et al.
Pubblicazione: (2025)
di: Xiang, Kunlan, et al.
Pubblicazione: (2025)
Systematic Categorization, Construction and Evaluation of New Attacks against Multi-modal Mobile GUI Agents
di: Yang, Yulong, et al.
Pubblicazione: (2024)
di: Yang, Yulong, et al.
Pubblicazione: (2024)
Towards Action Hijacking of Large Language Model-based Agent
di: Zhang, Yuyang, et al.
Pubblicazione: (2024)
di: Zhang, Yuyang, et al.
Pubblicazione: (2024)
Towards Physical World Backdoor Attacks against Skeleton Action Recognition
di: Zheng, Qichen, et al.
Pubblicazione: (2024)
di: Zheng, Qichen, et al.
Pubblicazione: (2024)
From ML to LLM: Evaluating the Robustness of Phishing Webpage Detection Models against Adversarial Attacks
di: Kulkarni, Aditya, et al.
Pubblicazione: (2024)
di: Kulkarni, Aditya, et al.
Pubblicazione: (2024)
The Trojan Example: Jailbreaking LLMs through Template Filling and Unsafety Reasoning
di: Liu, Mingrui, et al.
Pubblicazione: (2025)
di: Liu, Mingrui, et al.
Pubblicazione: (2025)
On the Adversarial Robustness of Large Vision-Language Models under Visual Token Compression
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
di: Zhang, Xinwei, et al.
Pubblicazione: (2026)
Adversarial Attack Based Countermeasures against Deep Learning Side-Channel Attacks
di: Gu, Ruizhe, et al.
Pubblicazione: (2020)
di: Gu, Ruizhe, et al.
Pubblicazione: (2020)
PentestGPT: An LLM-empowered Automatic Penetration Testing Tool
di: Deng, Gelei, et al.
Pubblicazione: (2023)
di: Deng, Gelei, et al.
Pubblicazione: (2023)
BURN: Backdoor Unlearning via Adversarial Boundary Analysis
di: Su, Yanghao, et al.
Pubblicazione: (2025)
di: Su, Yanghao, et al.
Pubblicazione: (2025)
Multi-task Adversarial Attacks against Black-box Model with Few-shot Queries
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
di: Wang, Wenqiang, et al.
Pubblicazione: (2025)
AutoEG: Exploiting Known Third-Party Vulnerabilities in Black-Box Web Applications
di: Yang, Ruozhao, et al.
Pubblicazione: (2026)
di: Yang, Ruozhao, et al.
Pubblicazione: (2026)
Capacitive Touchscreens at Risk: A Practical Side-Channel Attack on Smartphones via Electromagnetic Emanations
di: Cheng, Yukun, et al.
Pubblicazione: (2026)
di: Cheng, Yukun, et al.
Pubblicazione: (2026)
TEAM: Temporal Adversarial Examples Attack Model against Network Intrusion Detection System Applied to RNN
di: Liu, Ziyi, et al.
Pubblicazione: (2024)
di: Liu, Ziyi, et al.
Pubblicazione: (2024)
Supply-Chain Poisoning Attacks Against LLM Coding Agent Skill Ecosystems
di: Qu, Yubin, et al.
Pubblicazione: (2026)
di: Qu, Yubin, et al.
Pubblicazione: (2026)
CAAP: Capture-Aware Adversarial Patch Attacks on Palmprint Recognition Models
di: Liu, Renyang, et al.
Pubblicazione: (2026)
di: Liu, Renyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
When Search Goes Wrong: Red-Teaming Web-Augmented Large Language Models
di: Ou, Haoran, et al.
Pubblicazione: (2025) -
Oedipus: LLM-enchanced Reasoning CAPTCHA Solver
di: Deng, Gelei, et al.
Pubblicazione: (2024) -
Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
di: Zhao, Shiqian, et al.
Pubblicazione: (2025) -
SwitchPatch: Physical Adversarial Attack Strategy with Switchable Adversarial Objectives
di: Jiang, Hanrui, et al.
Pubblicazione: (2025) -
Turning Bias into Bugs: Bandit-Guided Style Manipulation Attacks on LLM Judges
di: Yang, Xianglin, et al.
Pubblicazione: (2026)