PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sehwag, Udari Madhushani, Shabihi, Shayan, McAvoy, Alex, Sehwag, Vikash, Xu, Yuancheng, Towers, Dalton, Huang, Furong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024)
In-Context Learning with Topological Information for Knowledge Graph Completion
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
Meeting times on graphs in near-cubic time
von: McAvoy, Alex
Veröffentlicht: (2026)
von: McAvoy, Alex
Veröffentlicht: (2026)
Can LLMs be Scammed? A Baseline Measurement Study
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
von: Xie, Tinghao, et al.
Veröffentlicht: (2024)
von: Xie, Tinghao, et al.
Veröffentlicht: (2024)
Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders
von: Campbell, David, et al.
Veröffentlicht: (2026)
von: Campbell, David, et al.
Veröffentlicht: (2026)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2026)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2026)
The coalescent in finite populations with arbitrary, fixed structure
von: Allen, Benjamin, et al.
Veröffentlicht: (2022)
von: Allen, Benjamin, et al.
Veröffentlicht: (2022)
Expectation-enforcing strategies for repeated games
von: Dimou, Nikos, et al.
Veröffentlicht: (2025)
von: Dimou, Nikos, et al.
Veröffentlicht: (2025)
Frequency-dependent returns in nonlinear public goods games
von: Hauert, Christoph, et al.
Veröffentlicht: (2024)
von: Hauert, Christoph, et al.
Veröffentlicht: (2024)
Theorizing to Cases: A Methodological Approach to Qualitative Normative Cases
von: Lauren Gatti, et al.
Veröffentlicht: (2024)
von: Lauren Gatti, et al.
Veröffentlicht: (2024)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
von: Cook, Thomas, et al.
Veröffentlicht: (2025)
von: Cook, Thomas, et al.
Veröffentlicht: (2025)
Edge Determining Sets and Determining Index
von: McAvoy, Sean, et al.
Veröffentlicht: (2022)
von: McAvoy, Sean, et al.
Veröffentlicht: (2022)
Promoting collective cooperation through temporal interactions
von: Meng, Yao, et al.
Veröffentlicht: (2024)
von: Meng, Yao, et al.
Veröffentlicht: (2024)
Behavioral alignment in social networks
von: Xia, Yu, et al.
Veröffentlicht: (2025)
von: Xia, Yu, et al.
Veröffentlicht: (2025)
Evaluating and Mitigating IP Infringement in Visual Generative AI
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2025)
von: Chakraborty, Souradip, et al.
Veröffentlicht: (2025)
AgentCrypt: Advancing Privacy and (Secure) Computation in AI Agent Collaboration
von: Karthikeyan, Harish, et al.
Veröffentlicht: (2025)
von: Karthikeyan, Harish, et al.
Veröffentlicht: (2025)
Symposium Introduction: Education for Democratic Sustainability and Transformation
von: Paula McAvoy, et al.
Veröffentlicht: (2024)
von: Paula McAvoy, et al.
Veröffentlicht: (2024)
How to Trace Latent Generative Model Generated Images without Artificial Watermark?
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
von: Wang, Zhenting, et al.
Veröffentlicht: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
von: Hayes, Kevin David, et al.
Veröffentlicht: (2025)
Finding needles in a haystack: A Black-Box Approach to Invisible Watermark Detection
von: Pan, Minzhou, et al.
Veröffentlicht: (2024)
von: Pan, Minzhou, et al.
Veröffentlicht: (2024)
Self-Comparison for Dataset-Level Membership Inference in Large (Vision-)Language Models
von: Ren, Jie, et al.
Veröffentlicht: (2024)
von: Ren, Jie, et al.
Veröffentlicht: (2024)
Adapting to Evolving Adversaries with Regularized Continual Robust Training
von: Dai, Sihui, et al.
Veröffentlicht: (2025)
von: Dai, Sihui, et al.
Veröffentlicht: (2025)
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI
von: Cheng, Siyuan, et al.
Veröffentlicht: (2025)
von: Cheng, Siyuan, et al.
Veröffentlicht: (2025)
A New Linear Scaling Rule for Private Adaptive Hyperparameter Optimization
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
von: Panda, Ashwinee, et al.
Veröffentlicht: (2022)
Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget
von: Sehwag, Vikash, et al.
Veröffentlicht: (2024)
von: Sehwag, Vikash, et al.
Veröffentlicht: (2024)
LHAW: Controllable Underspecification for Long-Horizon Tasks
von: Pu, George, et al.
Veröffentlicht: (2026)
von: Pu, George, et al.
Veröffentlicht: (2026)
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2026)
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2026)
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
von: Chao, Patrick, et al.
Veröffentlicht: (2024)
AegisLLM: Scaling Agentic Systems for Self-Reflective Defense in LLM Security
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
von: Cai, Zikui, et al.
Veröffentlicht: (2025)
Strategy evolution on temporal hypergraphs
von: Wang, Xiaochen, et al.
Veröffentlicht: (2023)
von: Wang, Xiaochen, et al.
Veröffentlicht: (2023)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
von: Wei, Boyi, et al.
Veröffentlicht: (2025)
von: Wei, Boyi, et al.
Veröffentlicht: (2025)
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
von: Akinfaderin, Adewale, et al.
Veröffentlicht: (2025)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
AI Risk Management Should Incorporate Both Safety and Security
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
von: Qi, Xiangyu, et al.
Veröffentlicht: (2024)
Intersection of Knowledge to Practice: A Purposeful Integration of Communication Partner Training in Aphasia With Adult Learning Principles for Healthcare Students
von: Catherine Torrington Eaton, et al.
Veröffentlicht: (2025)
von: Catherine Torrington Eaton, et al.
Veröffentlicht: (2025)
Does More Inference-Time Compute Really Help Robustness?
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Activity Recognition on Avatar-Anonymized Datasets with Masked Differential Privacy
von: Schneider, David, et al.
Veröffentlicht: (2024)
von: Schneider, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
von: Xu, Yuancheng, et al.
Veröffentlicht: (2024) -
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
von: Pathmanathan, Pankayaraj, et al.
Veröffentlicht: (2024) -
In-Context Learning with Topological Information for Knowledge Graph Completion
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024) -
Meeting times on graphs in near-cubic time
von: McAvoy, Alex
Veröffentlicht: (2026) -
Can LLMs be Scammed? A Baseline Measurement Study
von: Sehwag, Udari Madhushani, et al.
Veröffentlicht: (2024)