Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack
Fuente:
arXiv
Guardado en:
| Autores principales: | Fazla, Arnisa, Krauter, Lucas, Piedrahita, David Guzman, Michail, Andrianos |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
por: Michail, Andrianos, et al.
Publicado: (2025)
por: Michail, Andrianos, et al.
Publicado: (2025)
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
por: Przybyła, Piotr, et al.
Publicado: (2024)
por: Przybyła, Piotr, et al.
Publicado: (2024)
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
por: Moghe, Nikita, et al.
Publicado: (2024)
por: Moghe, Nikita, et al.
Publicado: (2024)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024)
por: Michail, Andrianos, et al.
Publicado: (2024)
Interpretable Text Embeddings and Text Similarity Explanation: A Survey
por: Opitz, Juri, et al.
Publicado: (2025)
por: Opitz, Juri, et al.
Publicado: (2025)
Adapting Multilingual Embedding Models to Historical Luxembourgish
por: Michail, Andrianos, et al.
Publicado: (2025)
por: Michail, Andrianos, et al.
Publicado: (2025)
Sentence Smith: Controllable Edits for Evaluating Text Embeddings
por: Li, Hongji, et al.
Publicado: (2025)
por: Li, Hongji, et al.
Publicado: (2025)
Evaluating Text Classification Robustness to Part-of-Speech Adversarial Examples
por: Samadi, Anahita, et al.
Publicado: (2024)
por: Samadi, Anahita, et al.
Publicado: (2024)
CAMOUFLAGE: Exploiting Misinformation Detection Systems Through LLM-driven Adversarial Claim Transformation
por: Bethany, Mazal, et al.
Publicado: (2025)
por: Bethany, Mazal, et al.
Publicado: (2025)
Information Representation Fairness in Long-Document Embeddings: The Peculiar Interaction of Positional and Language Bias
por: Schuhmacher, Elias, et al.
Publicado: (2026)
por: Schuhmacher, Elias, et al.
Publicado: (2026)
The Best Defense is Attack: Repairing Semantics in Textual Adversarial Examples
por: Yang, Heng, et al.
Publicado: (2023)
por: Yang, Heng, et al.
Publicado: (2023)
Not-in-Perspective: Towards Shielding Google's Perspective API Against Adversarial Negation Attacks
por: Alexiou, Michail S., et al.
Publicado: (2026)
por: Alexiou, Michail S., et al.
Publicado: (2026)
Arabic Synonym BERT-based Adversarial Examples for Text Classification
por: Alshahrani, Norah, et al.
Publicado: (2024)
por: Alshahrani, Norah, et al.
Publicado: (2024)
On the Robustness of Verbal Confidence of LLMs in Adversarial Attacks
por: Obadinma, Stephen, et al.
Publicado: (2025)
por: Obadinma, Stephen, et al.
Publicado: (2025)
Adversarial Attack for Explanation Robustness of Rationalization Models
por: Zhang, Yuankai, et al.
Publicado: (2024)
por: Zhang, Yuankai, et al.
Publicado: (2024)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
por: Koike, Ryuto, et al.
Publicado: (2023)
por: Koike, Ryuto, et al.
Publicado: (2023)
CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
por: Opitz, Juri, et al.
Publicado: (2026)
por: Opitz, Juri, et al.
Publicado: (2026)
Robustness of Large Language Models Against Adversarial Attacks
por: Tao, Yiyi, et al.
Publicado: (2024)
por: Tao, Yiyi, et al.
Publicado: (2024)
IAE: Irony-based Adversarial Examples for Sentiment Analysis Systems
por: Yi, Xiaoyin, et al.
Publicado: (2024)
por: Yi, Xiaoyin, et al.
Publicado: (2024)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
por: Aldahoul, Nouar, et al.
Publicado: (2025)
por: Aldahoul, Nouar, et al.
Publicado: (2025)
Robust Misinformation Detection by Visiting Potential Commonsense Conflict
por: Wang, Bing, et al.
Publicado: (2025)
por: Wang, Bing, et al.
Publicado: (2025)
DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising
por: Li, Zhenhao, et al.
Publicado: (2024)
por: Li, Zhenhao, et al.
Publicado: (2024)
A Classification-Guided Approach for Adversarial Attacks against Neural Machine Translation
por: Sadrizadeh, Sahar, et al.
Publicado: (2023)
por: Sadrizadeh, Sahar, et al.
Publicado: (2023)
Unpacking the Resilience of SNLI Contradiction Examples to Attacks
por: Verma, Chetan, et al.
Publicado: (2024)
por: Verma, Chetan, et al.
Publicado: (2024)
Camouflage is all you need: Evaluating and Enhancing Language Model Robustness Against Camouflage Adversarial Attacks
por: Huertas-García, Álvaro, et al.
Publicado: (2024)
por: Huertas-García, Álvaro, et al.
Publicado: (2024)
Exploring Adversarial Robustness in Classification tasks using DNA Language Models
por: Yoo, Hyunwoo, et al.
Publicado: (2024)
por: Yoo, Hyunwoo, et al.
Publicado: (2024)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
por: Piedrahita, David Guzman, et al.
Publicado: (2025)
por: Piedrahita, David Guzman, et al.
Publicado: (2025)
LANE: Lexical Adversarial Negative Examples for Word Sense Disambiguation
por: de Sá, Jader Martins Camboim, et al.
Publicado: (2025)
por: de Sá, Jader Martins Camboim, et al.
Publicado: (2025)
LLM Robustness Against Misinformation in Biomedical Question Answering
por: Bondarenko, Alexander, et al.
Publicado: (2024)
por: Bondarenko, Alexander, et al.
Publicado: (2024)
Fighting Fire with Fire: Adversarial Prompting to Generate a Misinformation Detection Dataset
por: Satapara, Shrey, et al.
Publicado: (2024)
por: Satapara, Shrey, et al.
Publicado: (2024)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
por: Sabir, Bushra, et al.
Publicado: (2023)
por: Sabir, Bushra, et al.
Publicado: (2023)
Robust Fake News Detection using Large Language Models under Adversarial Sentiment Attacks
por: Tahmasebi, Sahar, et al.
Publicado: (2026)
por: Tahmasebi, Sahar, et al.
Publicado: (2026)
Is LLM-as-a-Judge Robust? Investigating Universal Adversarial Attacks on Zero-shot LLM Assessment
por: Raina, Vyas, et al.
Publicado: (2024)
por: Raina, Vyas, et al.
Publicado: (2024)
On Adversarial Examples for Text Classification by Perturbing Latent Representations
por: Sooksatra, Korn, et al.
Publicado: (2024)
por: Sooksatra, Korn, et al.
Publicado: (2024)
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples
por: Yao, Jia-Yu, et al.
Publicado: (2023)
por: Yao, Jia-Yu, et al.
Publicado: (2023)
destroR: Attacking Transfer Models with Obfuscous Examples to Discard Perplexity
por: Ahmed, Saadat Rafid, et al.
Publicado: (2025)
por: Ahmed, Saadat Rafid, et al.
Publicado: (2025)
Goal-Aware Identification and Rectification of Misinformation in Multi-Agent Systems
por: Li, Zherui, et al.
Publicado: (2025)
por: Li, Zherui, et al.
Publicado: (2025)
Battling Misinformation: An Empirical Study on Adversarial Factuality in Open-Source Large Language Models
por: Sakib, Shahnewaz Karim, et al.
Publicado: (2025)
por: Sakib, Shahnewaz Karim, et al.
Publicado: (2025)
Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold Purification
por: Dang, Chenhao, et al.
Publicado: (2025)
por: Dang, Chenhao, et al.
Publicado: (2025)
Decoding Susceptibility: Modeling Misbelief to Misinformation Through a Computational Approach
por: Liu, Yanchen, et al.
Publicado: (2023)
por: Liu, Yanchen, et al.
Publicado: (2023)
Ejemplares similares
-
Examining Multilingual Embedding Models Cross-Lingually Through LLM-Generated Adversarial Examples
por: Michail, Andrianos, et al.
Publicado: (2025) -
Attacking Misinformation Detection Using Adversarial Examples Generated by Language Models
por: Przybyła, Piotr, et al.
Publicado: (2024) -
Machine Translation Meta Evaluation through Translation Accuracy Challenge Sets
por: Moghe, Nikita, et al.
Publicado: (2024) -
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024) -
Interpretable Text Embeddings and Text Similarity Explanation: A Survey
por: Opitz, Juri, et al.
Publicado: (2025)