A Constraint-Enforcing Reward for Adversarial Attacks on Text Classifiers
Fuente:
arXiv
Saved in:
| Main Authors: | Roth, Tom, Unanue, Inigo Jauregi, Abuadbba, Alsharif, Piccardi, Massimo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024)
by: Roth, Tom, et al.
Published: (2024)
SumTra: A Differentiable Pipeline for Few-Shot Cross-Lingual Summarization
by: Parnell, Jacob, et al.
Published: (2024)
by: Parnell, Jacob, et al.
Published: (2024)
Pref-CTRL: Preference Driven LLM Alignment using Representation Editing
by: Ashrafi, Imranul, et al.
Published: (2026)
by: Ashrafi, Imranul, et al.
Published: (2026)
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021)
by: Roth, Tom, et al.
Published: (2021)
ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark
by: Nguyen, Tung X., et al.
Published: (2026)
by: Nguyen, Tung X., et al.
Published: (2026)
Alert-ME: An Explainability-Driven Defense Against Adversarial Examples in Transformer-Based Text Classification
by: Sabir, Bushra, et al.
Published: (2023)
by: Sabir, Bushra, et al.
Published: (2023)
Adversarial Attacks Against Automated Fact-Checking: A Survey
by: Liu, Fanzhen, et al.
Published: (2025)
by: Liu, Fanzhen, et al.
Published: (2025)
DTO: a Differentiable Training Objective for Effective Counterfactual Story Rewriting
by: Girard, Amelia, et al.
Published: (2026)
by: Girard, Amelia, et al.
Published: (2026)
Mitigating Gradient Inversion Risks in Language Models via Token Obfuscation
by: Feng, Xinguo, et al.
Published: (2026)
by: Feng, Xinguo, et al.
Published: (2026)
An Investigation into Misuse of Java Security APIs by Large Language Models
by: Mousavi, Zahra, et al.
Published: (2024)
by: Mousavi, Zahra, et al.
Published: (2024)
Enforcing Temporal Constraints for LLM Agents
by: Kamath, Adharsh, et al.
Published: (2025)
by: Kamath, Adharsh, et al.
Published: (2025)
Large Language Model Adversarial Landscape Through the Lens of Attack Objectives
by: Wang, Nan, et al.
Published: (2025)
by: Wang, Nan, et al.
Published: (2025)
Your Extreme Multi-label Classifier is Secretly a Hierarchical Text Classifier for Free
by: Bertalis, Nerijus, et al.
Published: (2024)
by: Bertalis, Nerijus, et al.
Published: (2024)
Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text
by: Cheng, Yize, et al.
Published: (2025)
by: Cheng, Yize, et al.
Published: (2025)
Improving Vietnamese-English Medical Machine Translation
by: Vo, Nhu, et al.
Published: (2024)
by: Vo, Nhu, et al.
Published: (2024)
VertAttack: Taking advantage of Text Classifiers' horizontal vision
by: Rusert, Jonathan
Published: (2024)
by: Rusert, Jonathan
Published: (2024)
Automated Adversarial Discovery for Safety Classifiers
by: Lal, Yash Kumar, et al.
Published: (2024)
by: Lal, Yash Kumar, et al.
Published: (2024)
Multilingual LLM Prompting Strategies for Medical English-Vietnamese Machine Translation
by: Vo, Nhu, et al.
Published: (2025)
by: Vo, Nhu, et al.
Published: (2025)
Automatic Logical Forms improve fidelity in Table-to-Text generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
Towards Inclusive Toxic Content Moderation: Addressing Vulnerabilities to Adversarial Attacks in Toxicity Classifiers Tackling LLM-generated Content
by: Furniturewala, Shaz, et al.
Published: (2025)
by: Furniturewala, Shaz, et al.
Published: (2025)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
by: Liu, Han, et al.
Published: (2024)
by: Liu, Han, et al.
Published: (2024)
Adapters Mixup: Mixing Parameter-Efficient Adapters to Enhance the Adversarial Robustness of Fine-tuned Pre-trained Text Classifiers
by: Nguyen, Tuc, et al.
Published: (2024)
by: Nguyen, Tuc, et al.
Published: (2024)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
by: McKenzie, Ian R., et al.
Published: (2025)
by: McKenzie, Ian R., et al.
Published: (2025)
Vision-fused Attack: Advancing Aggressive and Stealthy Adversarial Text against Neural Machine Translation
by: Xue, Yanni, et al.
Published: (2024)
by: Xue, Yanni, et al.
Published: (2024)
PixT3: Pixel-based Table-To-Text Generation
by: Alonso, Iñigo, et al.
Published: (2023)
by: Alonso, Iñigo, et al.
Published: (2023)
Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modeling
by: Pathmanathan, Pankayaraj, et al.
Published: (2025)
by: Pathmanathan, Pankayaraj, et al.
Published: (2025)
Conflicts in Texts: Data, Implications and Challenges
by: Liu, Siyi, et al.
Published: (2025)
by: Liu, Siyi, et al.
Published: (2025)
From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
by: Birkenmaier, Lukas, et al.
Published: (2024)
by: Birkenmaier, Lukas, et al.
Published: (2024)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
by: Cheng, Myra, et al.
Published: (2024)
by: Cheng, Myra, et al.
Published: (2024)
Challenges in Explaining Pretrained Clinical Text Classifiers
by: Miok, Kristian, et al.
Published: (2026)
by: Miok, Kristian, et al.
Published: (2026)
A Modified Word Saliency-Based Adversarial Attack on Text Classification Models
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
Semantic Stealth: Adversarial Text Attacks on NLP Using Several Methods
by: Dey, Roopkatha, et al.
Published: (2024)
by: Dey, Roopkatha, et al.
Published: (2024)
Enhancing Adversarial Text Attacks on BERT Models with Projected Gradient Descent
by: Waghela, Hetvi, et al.
Published: (2024)
by: Waghela, Hetvi, et al.
Published: (2024)
Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations
by: Teja, Lekkala Sai, et al.
Published: (2025)
by: Teja, Lekkala Sai, et al.
Published: (2025)
The Hidden Cost of Modeling P(X): Vulnerability to Membership Inference Attacks in Generative Text Classifiers
by: Makroo, Owais, et al.
Published: (2025)
by: Makroo, Owais, et al.
Published: (2025)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
by: Li, Zelin, et al.
Published: (2024)
by: Li, Zelin, et al.
Published: (2024)
Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial Attacks
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Stance Detection: A Practical Guide to Classifying Political Beliefs in Text
by: Burnham, Michael
Published: (2023)
by: Burnham, Michael
Published: (2023)
OpenFact at CheckThat! 2024: Combining Multiple Attack Methods for Effective Adversarial Text Generation
by: Lewoniewski, Włodzimierz, et al.
Published: (2024)
by: Lewoniewski, Włodzimierz, et al.
Published: (2024)
ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks
by: Yu, Xiaodong, et al.
Published: (2023)
by: Yu, Xiaodong, et al.
Published: (2023)
Similar Items
-
A Generative Adversarial Attack for Multilingual Text Classifiers
by: Roth, Tom, et al.
Published: (2024) -
SumTra: A Differentiable Pipeline for Few-Shot Cross-Lingual Summarization
by: Parnell, Jacob, et al.
Published: (2024) -
Pref-CTRL: Preference Driven LLM Alignment using Representation Editing
by: Ashrafi, Imranul, et al.
Published: (2026) -
Token-Modification Adversarial Attacks for Natural Language Processing: A Survey
by: Roth, Tom, et al.
Published: (2021) -
ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & Benchmark
by: Nguyen, Tung X., et al.
Published: (2026)