Saved in:
| Main Authors: | Das, Nilanjana, Raff, Edward, Gaur, Manas |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.14644 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026)
by: Das, Nilanjana, et al.
Published: (2026)
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024)
by: Mohammadi, Seyedali, et al.
Published: (2024)
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024)
by: Mohseni, Seyedreza, et al.
Published: (2024)
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
by: Mohammadi, Seyedali, et al.
Published: (2026)
by: Mohammadi, Seyedali, et al.
Published: (2026)
Attribution in Scientific Literature: New Benchmark and Methods
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
Do LLMs Adhere to Label Definitions? Examining Their Receptivity to External Label Definitions
by: Mohammadi, Seyedali, et al.
Published: (2025)
by: Mohammadi, Seyedali, et al.
Published: (2025)
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
by: Bellibatlu, Rohith Reddy, et al.
Published: (2026)
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025)
by: Lamba, Naveen, et al.
Published: (2025)
SPRI: Aligning Large Language Models with Context-Situated Principles
by: Zhan, Hongli, et al.
Published: (2025)
by: Zhan, Hongli, et al.
Published: (2025)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
by: Gorti, Atmika, et al.
Published: (2024)
by: Gorti, Atmika, et al.
Published: (2024)
Neurosymbolic Retrievers for Retrieval-augmented Generation
by: Saxena, Yash, et al.
Published: (2026)
by: Saxena, Yash, et al.
Published: (2026)
TF-Attack: Transferable and Fast Adversarial Attacks on Large Language Models
by: Li, Zelin, et al.
Published: (2024)
by: Li, Zelin, et al.
Published: (2024)
Robustness of Large Language Models Against Adversarial Attacks
by: Tao, Yiyi, et al.
Published: (2024)
by: Tao, Yiyi, et al.
Published: (2024)
$\textit{LinkPrompt}$: Natural and Universal Adversarial Attacks on Prompt-based Language Models
by: Xu, Yue, et al.
Published: (2024)
by: Xu, Yue, et al.
Published: (2024)
SymLoc: Symbolic Localization of Hallucination across HaluEval and TruthfulQA
by: Lamba, Naveen, et al.
Published: (2025)
by: Lamba, Naveen, et al.
Published: (2025)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
by: Huang, Yukun, et al.
Published: (2024)
by: Huang, Yukun, et al.
Published: (2024)
Adversarial Confusion Attack: Disrupting Multimodal Large Language Models
by: Hoscilowicz, Jakub, et al.
Published: (2025)
by: Hoscilowicz, Jakub, et al.
Published: (2025)
Analyzing Chain of Thought (CoT) Approaches in Control Flow Code Deobfuscation Tasks
by: Mohseni, Seyedreza, et al.
Published: (2026)
by: Mohseni, Seyedreza, et al.
Published: (2026)
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Prompt Stealing Attacks Against Large Language Models
by: Sha, Zeyang, et al.
Published: (2024)
by: Sha, Zeyang, et al.
Published: (2024)
Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution
by: Saxena, Yash, et al.
Published: (2025)
by: Saxena, Yash, et al.
Published: (2025)
Skills-in-Context Prompting: Unlocking Compositionality in Large Language Models
by: Chen, Jiaao, et al.
Published: (2023)
by: Chen, Jiaao, et al.
Published: (2023)
Robustness of Prompting: Enhancing Robustness of Large Language Models Against Prompting Attacks
by: Mu, Lin, et al.
Published: (2025)
by: Mu, Lin, et al.
Published: (2025)
Adversarial Evasion Attack Efficiency against Large Language Models
by: Vitorino, João, et al.
Published: (2024)
by: Vitorino, João, et al.
Published: (2024)
Beyond Memorization: Testing LLM Reasoning on Unseen Theory of Computation Tasks
by: Shelat, Shlok, et al.
Published: (2026)
by: Shelat, Shlok, et al.
Published: (2026)
How Interpretable are Reasoning Explanations from Prompting Large Language Models?
by: Yeo, Wei Jie, et al.
Published: (2024)
by: Yeo, Wei Jie, et al.
Published: (2024)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
Membership Inference Attack against Long-Context Large Language Models
by: Wang, Zixiong, et al.
Published: (2024)
by: Wang, Zixiong, et al.
Published: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
by: Zhu, Kaijie, et al.
Published: (2023)
by: Zhu, Kaijie, et al.
Published: (2023)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024)
by: Govil, Priyanshul, et al.
Published: (2024)
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
by: Blau, Tsachi, et al.
Published: (2024)
by: Blau, Tsachi, et al.
Published: (2024)
Large Language Model based Situational Dialogues for Second Language Learning
by: Xu, Shuyao, et al.
Published: (2024)
by: Xu, Shuyao, et al.
Published: (2024)
Cultural Biases of Large Language Models and Humans in Historical Interpretation
by: Celli, Fabio, et al.
Published: (2025)
by: Celli, Fabio, et al.
Published: (2025)
ARREST: Adversarial Resilient Regulation Enhancing Safety and Truth in Large Language Models
by: Dasgupta, Sharanya, et al.
Published: (2026)
by: Dasgupta, Sharanya, et al.
Published: (2026)
The Resurgence of GCG Adversarial Attacks on Large Language Models
by: Tan, Yuting, et al.
Published: (2025)
by: Tan, Yuting, et al.
Published: (2025)
Context-aware Adversarial Attack on Named Entity Recognition
by: Chen, Shuguang, et al.
Published: (2023)
by: Chen, Shuguang, et al.
Published: (2023)
DROJ: A Prompt-Driven Attack against Large Language Models
by: Hu, Leyang, et al.
Published: (2024)
by: Hu, Leyang, et al.
Published: (2024)
Similar Items
-
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024) -
Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings
by: Das, Nilanjana, et al.
Published: (2026) -
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions
by: Mohammadi, Seyedali, et al.
Published: (2024) -
Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation
by: Mohseni, Seyedreza, et al.
Published: (2024) -
Experiments or Outcomes? Probing Scientific Feasibility in Large Language Models
by: Mohammadi, Seyedali, et al.
Published: (2026)