Universal Adversarial Triggers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arockiaraj, Benedict Florance, Feng, Alexander, Cai, Jianxiong, Cheng, Xiaoyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transfer Learning for Customized Car Racing Environments
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
Counting Machine Parts
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
Unraveling the Mystery of Scaling Laws: Part I
von: Su, Hui, et al.
Veröffentlicht: (2024)
von: Su, Hui, et al.
Veröffentlicht: (2024)
Rethinking LLM Memorization through the Lens of Adversarial Compression
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024)
WRAVAL -- WRiting Assist eVALuation
von: Benedict, Gabriel, et al.
Veröffentlicht: (2025)
von: Benedict, Gabriel, et al.
Veröffentlicht: (2025)
Self-playing Adversarial Language Game Enhances LLM Reasoning
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
von: Cheng, Pengyu, et al.
Veröffentlicht: (2024)
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
von: Liu, Qihao, et al.
Veröffentlicht: (2025)
Fine-Tuning Large Language Models to Appropriately Abstain with Semantic Entropy
von: Tjandra, Benedict Aaron, et al.
Veröffentlicht: (2024)
von: Tjandra, Benedict Aaron, et al.
Veröffentlicht: (2024)
Small Models are LLM Knowledge Triggers on Medical Tabular Prediction
von: Yan, Jiahuan, et al.
Veröffentlicht: (2024)
von: Yan, Jiahuan, et al.
Veröffentlicht: (2024)
SemRoDe: Macro Adversarial Training to Learn Representations That are Robust to Word-Level Attacks
von: Formento, Brian, et al.
Veröffentlicht: (2024)
von: Formento, Brian, et al.
Veröffentlicht: (2024)
MAR: Efficient Large Language Models via Module-aware Architecture Refinement
von: Cai, Junhong, et al.
Veröffentlicht: (2026)
von: Cai, Junhong, et al.
Veröffentlicht: (2026)
Linear Adversarial Concept Erasure
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
von: Ravfogel, Shauli, et al.
Veröffentlicht: (2022)
UT-ACA: Uncertainty-Triggered Adaptive Context Allocation for Long-Context Inference
von: Zhou, Lang, et al.
Veröffentlicht: (2026)
von: Zhou, Lang, et al.
Veröffentlicht: (2026)
Hidden Ads: Behavior Triggered Semantic Backdoors for Advertisement Injection in Vision Language Models
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
von: Yao, Duanyi, et al.
Veröffentlicht: (2026)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
von: Kraus, Oliver, et al.
Veröffentlicht: (2026)
von: Kraus, Oliver, et al.
Veröffentlicht: (2026)
Leveraging Open Information Extraction for More Robust Domain Transfer of Event Trigger Detection
von: Dukić, David, et al.
Veröffentlicht: (2023)
von: Dukić, David, et al.
Veröffentlicht: (2023)
Latent Adversarial Training Improves the Representation of Refusal
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
von: Abbas, Alexandra, et al.
Veröffentlicht: (2025)
Prompt Optimization via Adversarial In-Context Learning
von: Do, Xuan Long, et al.
Veröffentlicht: (2023)
von: Do, Xuan Long, et al.
Veröffentlicht: (2023)
Explaining the role of Intrinsic Dimensionality in Adversarial Training
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
von: Altinisik, Enes, et al.
Veröffentlicht: (2024)
Stochastic Adversarial Networks for Multi-Domain Text Classification
von: Wang, Xu, et al.
Veröffentlicht: (2024)
von: Wang, Xu, et al.
Veröffentlicht: (2024)
Adversarial Moment-Matching Distillation of Large Language Models
von: Jia, Chen
Veröffentlicht: (2024)
von: Jia, Chen
Veröffentlicht: (2024)
Adversarial Evasion Attack Efficiency against Large Language Models
von: Vitorino, João, et al.
Veröffentlicht: (2024)
von: Vitorino, João, et al.
Veröffentlicht: (2024)
Assessing Adversarial Robustness of Large Language Models: An Empirical Study
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
von: Yang, Zeyu, et al.
Veröffentlicht: (2024)
Evaluating Text Classification Robustness to Part-of-Speech Adversarial Examples
von: Samadi, Anahita, et al.
Veröffentlicht: (2024)
von: Samadi, Anahita, et al.
Veröffentlicht: (2024)
Self-Refining Language Model Anonymizers via Adversarial Distillation
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2025)
von: Kim, Kyuyoung, et al.
Veröffentlicht: (2025)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
von: Zhang, Collin, et al.
Veröffentlicht: (2024)
Adversarial Demonstration Learning for Low-resource NER Using Dual Similarity
von: Yuan, Guowen, et al.
Veröffentlicht: (2025)
von: Yuan, Guowen, et al.
Veröffentlicht: (2025)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
Margin Discrepancy-based Adversarial Training for Multi-Domain Text Classification
von: Wu, Yuan
Veröffentlicht: (2024)
von: Wu, Yuan
Veröffentlicht: (2024)
Multi-Trigger Poisoning Amplifies Backdoor Vulnerabilities in LLMs
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
von: Sivapiromrat, Sanhanat, et al.
Veröffentlicht: (2025)
Adversarial Tokenization
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
von: Geh, Renato Lui, et al.
Veröffentlicht: (2025)
AR$^2$: Adversarial Reinforcement Learning for Abstract Reasoning in Large Language Models
von: Yeh, Cheng-Kai, et al.
Veröffentlicht: (2025)
von: Yeh, Cheng-Kai, et al.
Veröffentlicht: (2025)
Token-Level Adversarial Prompt Detection Based on Perplexity Measures and Contextual Information
von: Hu, Zhengmian, et al.
Veröffentlicht: (2023)
von: Hu, Zhengmian, et al.
Veröffentlicht: (2023)
PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning
von: Zhang, Tianrong, et al.
Veröffentlicht: (2024)
von: Zhang, Tianrong, et al.
Veröffentlicht: (2024)
DACTYL: Diverse Adversarial Corpus of Texts Yielded from Large Language Models
von: Thorat, Shantanu, et al.
Veröffentlicht: (2025)
von: Thorat, Shantanu, et al.
Veröffentlicht: (2025)
From Insight to Exploit: Leveraging LLM Collaboration for Adaptive Adversarial Text Generation
von: Sultana, Najrin, et al.
Veröffentlicht: (2025)
von: Sultana, Najrin, et al.
Veröffentlicht: (2025)
CAARMA: Class Augmentation with Adversarial Mixup Regularization
von: Baali, Massa, et al.
Veröffentlicht: (2025)
von: Baali, Massa, et al.
Veröffentlicht: (2025)
Unsupervised Text Embedding Space Generation Using Generative Adversarial Networks for Text Synthesis
von: Lee, Jun-Min, et al.
Veröffentlicht: (2023)
von: Lee, Jun-Min, et al.
Veröffentlicht: (2023)
Future Events as Backdoor Triggers: Investigating Temporal Vulnerabilities in LLMs
von: Price, Sara, et al.
Veröffentlicht: (2024)
von: Price, Sara, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transfer Learning for Customized Car Racing Environments
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026) -
Counting Machine Parts
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026) -
Unraveling the Mystery of Scaling Laws: Part I
von: Su, Hui, et al.
Veröffentlicht: (2024) -
Rethinking LLM Memorization through the Lens of Adversarial Compression
von: Schwarzschild, Avi, et al.
Veröffentlicht: (2024) -
WRAVAL -- WRiting Assist eVALuation
von: Benedict, Gabriel, et al.
Veröffentlicht: (2025)