Investigating Adversarial Trigger Transfer in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Meade, Nicholas, Patel, Arkil, Reddy, Siva |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2025)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
von: Adlakha, Vaibhav, et al.
Veröffentlicht: (2023)
von: Adlakha, Vaibhav, et al.
Veröffentlicht: (2023)
Scope Ambiguities in Large Language Models
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
von: Kamath, Gaurav, et al.
Veröffentlicht: (2024)
Are self-explanations from Large Language Models faithful?
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
von: Madsen, Andreas, et al.
Veröffentlicht: (2024)
Language Models Largely Exhibit Human-like Constituent Ordering Preferences
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)
DeepSeek-R1 Thoughtology: Let's think about LLM Reasoning
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
von: Marjanović, Sara Vera, et al.
Veröffentlicht: (2025)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
AutoAdv: Automated Adversarial Prompting for Multi-Turn Jailbreaking of Large Language Models
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
von: Reddy, Aashray, et al.
Veröffentlicht: (2025)
LLM2Vec-Gen: Generative Embeddings from Large Language Models
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2026)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2026)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
von: Li, Zelin, et al.
Veröffentlicht: (2024)
von: Li, Zelin, et al.
Veröffentlicht: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
von: Ball, Sarah, et al.
Veröffentlicht: (2025)
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2024)
BELL: Benchmarking the Explainability of Large Language Models
von: Ahmed, Syed Quiser, et al.
Veröffentlicht: (2025)
von: Ahmed, Syed Quiser, et al.
Veröffentlicht: (2025)
Not All Data Are Unlearned Equally
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
von: Krishnan, Aravind, et al.
Veröffentlicht: (2025)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
von: Liu, Hongfu, et al.
Veröffentlicht: (2024)
Universal Adversarial Triggers
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
von: Arockiaraj, Benedict Florance, et al.
Veröffentlicht: (2026)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
When does word order matter and when doesn't it?
von: Chen, Xuanda, et al.
Veröffentlicht: (2024)
von: Chen, Xuanda, et al.
Veröffentlicht: (2024)
If there's a Trigger Warning, then where's the Trigger? Investigating Trigger Warnings at the Passage Level
von: Wiegmann, Matti, et al.
Veröffentlicht: (2024)
von: Wiegmann, Matti, et al.
Veröffentlicht: (2024)
Investigating Political and Demographic Associations in Large Language Models Through Moral Foundations Theory
von: Smith-Vaniz, Nicole, et al.
Veröffentlicht: (2025)
von: Smith-Vaniz, Nicole, et al.
Veröffentlicht: (2025)
Analyzing the Safety of Japanese Large Language Models in Stereotype-Triggering Prompts
von: Nakanishi, Akito, et al.
Veröffentlicht: (2025)
von: Nakanishi, Akito, et al.
Veröffentlicht: (2025)
Robustness of Large Language Models Against Adversarial Attacks
von: Tao, Yiyi, et al.
Veröffentlicht: (2024)
von: Tao, Yiyi, et al.
Veröffentlicht: (2024)
A Compositional Typed Semantics for Universal Dependencies
von: Bradford, Laurestine, et al.
Veröffentlicht: (2024)
von: Bradford, Laurestine, et al.
Veröffentlicht: (2024)
The StatCan Dialogue Dataset: Retrieving Data Tables through Conversations with Genuine Intents
von: Lu, Xing Han, et al.
Veröffentlicht: (2023)
von: Lu, Xing Han, et al.
Veröffentlicht: (2023)
Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
von: Lasnier, Théo, et al.
Veröffentlicht: (2026)
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models
von: Rajeev, Meghana, et al.
Veröffentlicht: (2025)
von: Rajeev, Meghana, et al.
Veröffentlicht: (2025)
Investigating Numerical Translation with Large Language Models
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
Investigating Model Editing for Unlearning in Large Language Models
von: Hossain, Shariqah, et al.
Veröffentlicht: (2025)
von: Hossain, Shariqah, et al.
Veröffentlicht: (2025)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
von: Patel, Ajay, et al.
Veröffentlicht: (2022)
Are Large Language Models Truly Smarter Than Humans?
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
von: M, Eshwar Reddy, et al.
Veröffentlicht: (2026)
Investigating Cultural Alignment of Large Language Models
von: AlKhamissi, Badr, et al.
Veröffentlicht: (2024)
von: AlKhamissi, Badr, et al.
Veröffentlicht: (2024)
AdvSumm: Adversarial Training for Bias Mitigation in Text Summarization
von: Gupta, Mukur, et al.
Veröffentlicht: (2025)
von: Gupta, Mukur, et al.
Veröffentlicht: (2025)
Cross-Lingual Optimization for Language Transfer in Large Language Models
von: Lee, Jungseob, et al.
Veröffentlicht: (2025)
von: Lee, Jungseob, et al.
Veröffentlicht: (2025)
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
von: Hoscilowicz, Jakub, et al.
Veröffentlicht: (2025)
von: Hoscilowicz, Jakub, et al.
Veröffentlicht: (2025)
On Adversarial Robustness of Language Models in Transfer Learning
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
von: Turbal, Bohdan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025) -
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026) -
Exploiting Instruction-Following Retrievers for Malicious Information Retrieval
von: BehnamGhader, Parishad, et al.
Veröffentlicht: (2025) -
SafeArena: Evaluating the Safety of Autonomous Web Agents
von: Tur, Ada Defne, et al.
Veröffentlicht: (2025)