Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Tingchen, Barez, Fazl |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
Deceiving Question-Answering Models: A Hybrid Word-Level Adversarial Approach
von: Li, Jiyao, et al.
Veröffentlicht: (2024)
von: Li, Jiyao, et al.
Veröffentlicht: (2024)
Best-of-N Jailbreaking
von: Hughes, John, et al.
Veröffentlicht: (2024)
von: Hughes, John, et al.
Veröffentlicht: (2024)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
von: Sheshadri, Abhay, et al.
Veröffentlicht: (2024)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Augmenting Math Word Problems via Iterative Question Composing
von: Liu, Haoxiong, et al.
Veröffentlicht: (2024)
von: Liu, Haoxiong, et al.
Veröffentlicht: (2024)
Beyond Prompting: An Efficient Embedding Framework for Open-Domain Question Answering
von: Hu, Zhanghao, et al.
Veröffentlicht: (2025)
von: Hu, Zhanghao, et al.
Veröffentlicht: (2025)
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
von: Wang, Tony T., et al.
Veröffentlicht: (2024)
Words as Beacons: Guiding RL Agents with High-Level Language Prompts
von: Ruiz-Gonzalez, Unai, et al.
Veröffentlicht: (2024)
von: Ruiz-Gonzalez, Unai, et al.
Veröffentlicht: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
We're Different, We're the Same: Creative Homogeneity Across LLMs
von: Wenger, Emily, et al.
Veröffentlicht: (2025)
von: Wenger, Emily, et al.
Veröffentlicht: (2025)
Chain-of-Thought Hijacking
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
von: Zhao, Jianli, et al.
Veröffentlicht: (2025)
Guided Perturbation Sensitivity (GPS): Detecting Adversarial Text via Embedding Stability and Word Importance
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
von: Tuck, Bryan E., et al.
Veröffentlicht: (2025)
Rainbow Teaming: Open-Ended Generation of Diverse Adversarial Prompts
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
von: Samvelyan, Mikayel, et al.
Veröffentlicht: (2024)
Word-Sequence Entropy: Towards Uncertainty Estimation in Free-Form Medical Question Answering Applications and Beyond
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyuan, et al.
Veröffentlicht: (2024)
Scalable Prompt Routing via Fine-Grained Latent Task Discovery
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
von: Zhang, Yunyi, et al.
Veröffentlicht: (2026)
PromptWizard: Task-Aware Prompt Optimization Framework
von: Agarwal, Eshaan, et al.
Veröffentlicht: (2024)
von: Agarwal, Eshaan, et al.
Veröffentlicht: (2024)
What's the Magic Word? A Control Theory of LLM Prompting
von: Bhargava, Aman, et al.
Veröffentlicht: (2023)
von: Bhargava, Aman, et al.
Veröffentlicht: (2023)
Enhancing NLP Robustness and Generalization through LLM-Generated Contrast Sets: A Scalable Framework for Systematic Evaluation and Adversarial Training
von: Lin, Hender
Veröffentlicht: (2025)
von: Lin, Hender
Veröffentlicht: (2025)
Evaluating Adversarial Robustness of Concept Representations in Sparse Autoencoders
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
von: Li, Aaron J., et al.
Veröffentlicht: (2025)
Query-Based Adversarial Prompt Generation
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
RobustSentEmbed: Robust Sentence Embeddings Using Adversarial Self-Supervised Contrastive Learning
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
von: Asl, Javad Rafiei, et al.
Veröffentlicht: (2024)
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
von: Neitemeier, Pit, et al.
Veröffentlicht: (2025)
von: Neitemeier, Pit, et al.
Veröffentlicht: (2025)
Ask more, know better: Reinforce-Learned Prompt Questions for Decision Making with Large Language Models
von: Yan, Xue, et al.
Veröffentlicht: (2023)
von: Yan, Xue, et al.
Veröffentlicht: (2023)
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
von: Li, Yubo, et al.
Veröffentlicht: (2025)
von: Li, Yubo, et al.
Veröffentlicht: (2025)
A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
von: Wang, Zhanliang, et al.
Veröffentlicht: (2026)
MultiQ&A: An Analysis in Measuring Robustness via Automated Crowdsourcing of Question Perturbations and Answers
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
von: Cho, Nicole, et al.
Veröffentlicht: (2025)
Bridging Robustness and Generalization Against Word Substitution Attacks in NLP via the Growth Bound Matrix Approach
von: Bouri, Mohammed, et al.
Veröffentlicht: (2025)
von: Bouri, Mohammed, et al.
Veröffentlicht: (2025)
Speaking the Same Language: Leveraging LLMs in Standardizing Clinical Data for AI
von: Sett, Arindam, et al.
Veröffentlicht: (2024)
von: Sett, Arindam, et al.
Veröffentlicht: (2024)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language Models
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
von: Yuan, Hongbang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023) -
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025) -
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)