Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Przybyła, Piotr, McGill, Euan, Saggion, Horacio |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sign Language Gloss Embedding Models
by: McGill, Euan, et al.
Published: (2024)
by: McGill, Euan, et al.
Published: (2024)
Verifying the Robustness of Automatic Credibility Assessment
by: Przybyła, Piotr, et al.
Published: (2023)
by: Przybyła, Piotr, et al.
Published: (2023)
Deanthropomorphising NLP: Can a Language Model Be Conscious?
by: Shardlow, Matthew, et al.
Published: (2022)
by: Shardlow, Matthew, et al.
Published: (2022)
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
by: Fazla, Arnisa, et al.
Published: (2025)
by: Fazla, Arnisa, et al.
Published: (2025)
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)
by: Rybak, Piotr, et al.
Published: (2022)
Towards Trustworthy Lexical Simplification: Exploring Safety and Efficiency with Small LLMs
by: Hayakawa, Akio, et al.
Published: (2025)
by: Hayakawa, Akio, et al.
Published: (2025)
The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples
by: Yang, Heng, et al.
Published: (2023)
by: Yang, Heng, et al.
Published: (2023)
A Multilingual Human Annotated Corpus of Original and Easy-to-Read Texts to Support Access to Democratic Participatory Processes
by: Bott, Stefan, et al.
Published: (2026)
by: Bott, Stefan, et al.
Published: (2026)
Fighting Fire with Fire: Adversarial Prompting to Generate a Misinformation Detection Dataset
by: Satapara, Shrey, et al.
Published: (2024)
by: Satapara, Shrey, et al.
Published: (2024)
Simple is not Enough: Document-level Text Simplification using Readability and Coherence
by: Vásquez-Rodríguez, Laura, et al.
Published: (2024)
by: Vásquez-Rodríguez, Laura, et al.
Published: (2024)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
by: Koike, Ryuto, et al.
Published: (2023)
by: Koike, Ryuto, et al.
Published: (2023)
FMDLlama: Financial Misinformation Detection based on Large Language Models
by: Liu, Zhiwei, et al.
Published: (2024)
by: Liu, Zhiwei, et al.
Published: (2024)
Misinformation Detection using Large Language Models with Explainability
by: Patel, Jainee, et al.
Published: (2025)
by: Patel, Jainee, et al.
Published: (2025)
Lexical Complexity Prediction and Lexical Simplification for Catalan and Spanish: Resource Creation, Quality Assessment, and Ethical Considerations
by: Bott, Stefan, et al.
Published: (2024)
by: Bott, Stefan, et al.
Published: (2024)
Testing the Generalization of Neural Language Models for COVID-19 Misinformation Detection
by: Wahle, Jan Philip, et al.
Published: (2021)
by: Wahle, Jan Philip, et al.
Published: (2021)
ReviewScore: Misinformed Peer Review Detection with Large Language Models
by: Ryu, Hyun, et al.
Published: (2025)
by: Ryu, Hyun, et al.
Published: (2025)
Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models
by: Sakib, Shahnewaz Karim, et al.
Published: (2025)
by: Sakib, Shahnewaz Karim, et al.
Published: (2025)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
by: Li, Zelin, et al.
Published: (2024)
by: Li, Zelin, et al.
Published: (2024)
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
by: Bailey, Luke, et al.
Published: (2023)
by: Bailey, Luke, et al.
Published: (2023)
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
by: Bethany, Mazal, et al.
Published: (2025)
by: Bethany, Mazal, et al.
Published: (2025)
Unified Large Language Models for Misinformation Detection in Low-Resource Linguistic Settings
by: Islam, Muhammad, et al.
Published: (2025)
by: Islam, Muhammad, et al.
Published: (2025)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks
by: Tahmasebi, Sahar, et al.
Published: (2026)
by: Tahmasebi, Sahar, et al.
Published: (2026)
Multimodal Misinformation Detection using Large Vision-Language Models
by: Tahmasebi, Sahar, et al.
Published: (2024)
by: Tahmasebi, Sahar, et al.
Published: (2024)
Unlearning Climate Misinformation in Large Language Models
by: Fore, Michael, et al.
Published: (2024)
by: Fore, Michael, et al.
Published: (2024)
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
by: Michail, Andrianos, et al.
Published: (2025)
by: Michail, Andrianos, et al.
Published: (2025)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
by: Hoscilowicz, Jakub, et al.
Published: (2025)
by: Hoscilowicz, Jakub, et al.
Published: (2025)
Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
by: Ren, Juan, et al.
Published: (2025)
by: Ren, Juan, et al.
Published: (2025)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
by: Teja, Lekkala Sai, et al.
Published: (2025)
by: Teja, Lekkala Sai, et al.
Published: (2025)
Images Amplify Misinformation Sharing in Vision-Language Models
by: Plebe, Alice, et al.
Published: (2025)
by: Plebe, Alice, et al.
Published: (2025)
HiEAG: Evidence-Augmented Generation for Out-of-Context Misinformation Detection
by: Wu, Junjie, et al.
Published: (2025)
by: Wu, Junjie, et al.
Published: (2025)
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models
by: Han, Chen, et al.
Published: (2025)
by: Han, Chen, et al.
Published: (2025)
Exploring Gradient-Guided Masked Language Model to Detect Textual Adversarial Attacks
by: Zhang, Xiaomei, et al.
Published: (2025)
by: Zhang, Xiaomei, et al.
Published: (2025)
Emotion Detection for Misinformation: A Review
by: Liu, Zhiwei, et al.
Published: (2023)
by: Liu, Zhiwei, et al.
Published: (2023)
Adversarial Attacks on AI-Generated Text Detection Models: A Token Probability-Based Approach Using Embeddings
by: Kadhim, Ahmed K., et al.
Published: (2025)
by: Kadhim, Ahmed K., et al.
Published: (2025)
Generative Debunking of Climate Misinformation
by: Zanartu, Francisco, et al.
Published: (2024)
by: Zanartu, Francisco, et al.
Published: (2024)
Monolingual and Multilingual Misinformation Detection for Low-Resource Languages: A Comprehensive Survey
by: Wang, Xinyu, et al.
Published: (2024)
by: Wang, Xinyu, et al.
Published: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
by: Aldahoul, Nouar, et al.
Published: (2025)
by: Aldahoul, Nouar, et al.
Published: (2025)
Can LLM-Generated Misinformation Be Detected?
by: Chen, Canyu, et al.
Published: (2023)
by: Chen, Canyu, et al.
Published: (2023)
Similar Items
-
Sign Language Gloss Embedding Models
by: McGill, Euan, et al.
Published: (2024) -
Verifying the Robustness of Automatic Credibility Assessment
by: Przybyła, Piotr, et al.
Published: (2023) -
Deanthropomorphising NLP: Can a Language Model Be Conscious?
by: Shardlow, Matthew, et al.
Published: (2022) -
Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
by: Fazla, Arnisa, et al.
Published: (2025) -
PolQA: Polish Question Answering Dataset
by: Rybak, Piotr, et al.
Published: (2022)