Enhancing LLM Agent Safety via Causal Influence Prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Hahm, Dongyoon, Jin, Woogyeol, Choi, June Suk, Ahn, Sungsoo, Lee, Kimin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
by: Hahm, Dongyoon, et al.
Published: (2025)
by: Hahm, Dongyoon, et al.
Published: (2025)
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026)
by: Hahm, Dongyoon, et al.
Published: (2026)
Benchmarking Mobile Device Control Agents across Diverse Configurations
by: Lee, Juyong, et al.
Published: (2024)
by: Lee, Juyong, et al.
Published: (2024)
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)
by: Kim, Dongyoung, et al.
Published: (2024)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
by: Yoon, Hyungjun, et al.
Published: (2024)
by: Yoon, Hyungjun, et al.
Published: (2024)
Lost in the Prompt Order: Revealing the Limitations of Causal Attention in Language Models
by: Ok, Hyunjong, et al.
Published: (2026)
by: Ok, Hyunjong, et al.
Published: (2026)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
by: Jin, Kyohoon, et al.
Published: (2024)
by: Jin, Kyohoon, et al.
Published: (2024)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
by: Kim, Taesu, et al.
Published: (2024)
by: Kim, Taesu, et al.
Published: (2024)
Certifying LLM Safety against Adversarial Prompting
by: Kumar, Aounon, et al.
Published: (2023)
by: Kumar, Aounon, et al.
Published: (2023)
Understanding Impact of Human Feedback via Influence Functions
by: Min, Taywon, et al.
Published: (2025)
by: Min, Taywon, et al.
Published: (2025)
LLM4Causal: Democratized Causal Tools for Everyone via Large Language Model
by: Jiang, Haitao, et al.
Published: (2023)
by: Jiang, Haitao, et al.
Published: (2023)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
by: Hadeliya, Tsimur, et al.
Published: (2025)
by: Hadeliya, Tsimur, et al.
Published: (2025)
StateAct: Enhancing LLM Base Agents via Self-prompting and State-tracking
by: Rozanov, Nikolai, et al.
Published: (2024)
by: Rozanov, Nikolai, et al.
Published: (2024)
RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
How Susceptible are LLMs to Influence in Prompts?
by: Anagnostidis, Sotiris, et al.
Published: (2024)
by: Anagnostidis, Sotiris, et al.
Published: (2024)
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
by: Lingo, Ryan, et al.
Published: (2024)
by: Lingo, Ryan, et al.
Published: (2024)
Structural Reasoning Improves Molecular Understanding of LLM
by: Jang, Yunhui, et al.
Published: (2024)
by: Jang, Yunhui, et al.
Published: (2024)
Prompting Fairness: Integrating Causality to Debias Large Language Models
by: Li, Jingling, et al.
Published: (2024)
by: Li, Jingling, et al.
Published: (2024)
No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping
by: Le, Thanh-Long V., et al.
Published: (2025)
by: Le, Thanh-Long V., et al.
Published: (2025)
ExpeL: LLM Agents Are Experiential Learners
by: Zhao, Andrew, et al.
Published: (2023)
by: Zhao, Andrew, et al.
Published: (2023)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection
by: Kim, Jeonghye, et al.
Published: (2025)
by: Kim, Jeonghye, et al.
Published: (2025)
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
by: Choe, Sang Keun, et al.
Published: (2024)
by: Choe, Sang Keun, et al.
Published: (2024)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
LongSafety: Enhance Safety for Long-Context LLMs
by: Huang, Mianqiu, et al.
Published: (2024)
by: Huang, Mianqiu, et al.
Published: (2024)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
by: Yin, Yueqin, et al.
Published: (2024)
by: Yin, Yueqin, et al.
Published: (2024)
BAGEN: Are LLM Agents Budget-Aware?
by: Lin, Yuxiang, et al.
Published: (2026)
by: Lin, Yuxiang, et al.
Published: (2026)
System Prompt Optimization with Meta-Learning
by: Choi, Yumin, et al.
Published: (2025)
by: Choi, Yumin, et al.
Published: (2025)
LLMScan: Causal Scan for LLM Misbehavior Detection
by: Zhang, Mengdi, et al.
Published: (2024)
by: Zhang, Mengdi, et al.
Published: (2024)
RePrompt: Planning by Automatic Prompt Engineering for Large Language Models Agents
by: Chen, Weizhe, et al.
Published: (2024)
by: Chen, Weizhe, et al.
Published: (2024)
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
by: Shen, Han, et al.
Published: (2024)
by: Shen, Han, et al.
Published: (2024)
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
by: Jin, Jikai, et al.
Published: (2025)
by: Jin, Jikai, et al.
Published: (2025)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
by: Ahn, Jihyun Janice, et al.
Published: (2024)
by: Ahn, Jihyun Janice, et al.
Published: (2024)
IDIAPers @ Causal News Corpus 2022: Efficient Causal Relation Identification Through a Prompt-based Few-shot Approach
by: Burdisso, Sergio, et al.
Published: (2022)
by: Burdisso, Sergio, et al.
Published: (2022)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
by: Huang, Wei, et al.
Published: (2026)
by: Huang, Wei, et al.
Published: (2026)
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization
by: Fu, Lucheng, et al.
Published: (2026)
by: Fu, Lucheng, et al.
Published: (2026)
TrajAgent: An LLM-Agent Framework for Trajectory Modeling via Large-and-Small Model Collaboration
by: Du, Yuwei, et al.
Published: (2024)
by: Du, Yuwei, et al.
Published: (2024)
Causally-Enhanced Reinforcement Policy Optimization
by: Wang, Xiangqi, et al.
Published: (2025)
by: Wang, Xiangqi, et al.
Published: (2025)
Evaluating Very Long-Term Conversational Memory of LLM Agents
by: Maharana, Adyasha, et al.
Published: (2024)
by: Maharana, Adyasha, et al.
Published: (2024)
Similar Items
-
Unintended Misalignment from Agentic Fine-Tuning: Risks and Mitigation
by: Hahm, Dongyoon, et al.
Published: (2025) -
MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device Control
by: Lee, Juyong, et al.
Published: (2024) -
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
by: Hahm, Dongyoon, et al.
Published: (2026) -
Benchmarking Mobile Device Control Agents across Diverse Configurations
by: Lee, Juyong, et al.
Published: (2024) -
Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
by: Kim, Dongyoung, et al.
Published: (2024)