Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Bonagiri, Vamshi Krishna, Kumaragurum, Ponnurangam, Nguyen, Khanh, Plaut, Benjamin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026)
by: Plaut, Benjamin
Published: (2026)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024)
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
by: Joishy, Anish R, et al.
Published: (2025)
by: Joishy, Anish R, et al.
Published: (2025)
COBIAS: Assessing the Contextual Reliability of Bias Benchmarks for Language Models
by: Govil, Priyanshul, et al.
Published: (2024)
by: Govil, Priyanshul, et al.
Published: (2024)
Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
by: Plaut, Benjamin, et al.
Published: (2024)
by: Plaut, Benjamin, et al.
Published: (2024)
From Human Judgements to Predictive Models: Unravelling Acceptability in Code-Mixed Sentences
by: Kodali, Prashant, et al.
Published: (2024)
by: Kodali, Prashant, et al.
Published: (2024)
Cultivating Game Sense for Yourself: Making VLMs Gaming Experts
by: Lu, Wenxuan, et al.
Published: (2025)
by: Lu, Wenxuan, et al.
Published: (2025)
Train Yourself as an LLM: Exploring Effects of AI Literacy on Persuasion via Role-playing LLM Training
by: Fan, Qihui, et al.
Published: (2026)
by: Fan, Qihui, et al.
Published: (2026)
Talking to Yourself: Defying Forgetting in Large Language Models
by: Sun, Yutao, et al.
Published: (2026)
by: Sun, Yutao, et al.
Published: (2026)
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation
by: Guo, Ruohao, et al.
Published: (2025)
by: Guo, Ruohao, et al.
Published: (2025)
Keeping Yourself is Important in Downstream Tuning Multimodal Large Language Model
by: Huang, Wenke, et al.
Published: (2025)
by: Huang, Wenke, et al.
Published: (2025)
Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning
by: Huang, Wenke, et al.
Published: (2024)
by: Huang, Wenke, et al.
Published: (2024)
Titanic Calling: Low Bandwidth Video Conference from the Titanic Wreck
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
by: Eyiokur, Fevziye Irem, et al.
Published: (2024)
Look Before You Leap: Autonomous Exploration for LLM Agents
by: Ye, Ziang, et al.
Published: (2026)
by: Ye, Ziang, et al.
Published: (2026)
You Have No One to Blame But Yourself.
by: Vivrette, Lyndon, et al.
Published: (1984)
by: Vivrette, Lyndon, et al.
Published: (1984)
UP2You: Fast Reconstruction of Yourself from Unconstrained Photo Collections
by: Cai, Zeyu, et al.
Published: (2025)
by: Cai, Zeyu, et al.
Published: (2025)
Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing
by: Yuan, Wenhao, et al.
Published: (2026)
by: Yuan, Wenhao, et al.
Published: (2026)
Manage Yourself.
by: Wall, Barbara
Published: (2003)
by: Wall, Barbara
Published: (2003)
ViFactCheck: A New Benchmark Dataset and Methods for Multi-domain News Fact-Checking in Vietnamese
by: Hoa, Tran Thai, et al.
Published: (2024)
by: Hoa, Tran Thai, et al.
Published: (2024)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
by: Luo, Jiani, et al.
Published: (2026)
by: Luo, Jiani, et al.
Published: (2026)
Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents
by: Zeng, Jingying, et al.
Published: (2025)
by: Zeng, Jingying, et al.
Published: (2025)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
SnapKV: LLM Knows What You are Looking for Before Generation
by: Li, Yuhong, et al.
Published: (2024)
by: Li, Yuhong, et al.
Published: (2024)
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA
by: Dong, Xuanzhao, et al.
Published: (2025)
by: Dong, Xuanzhao, et al.
Published: (2025)
Assessing the Quality and Safety of Do‐It‐Yourself Cosmetic Neuromodulator Injection Tutorials on YouTube
by: Lauren Ching, et al.
Published: (2025)
by: Lauren Ching, et al.
Published: (2025)
Agent-SafetyBench: Evaluating the Safety of LLM Agents
by: Zhang, Zhexin, et al.
Published: (2024)
by: Zhang, Zhexin, et al.
Published: (2024)
Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
by: Xiang, Yang, et al.
Published: (2025)
by: Xiang, Yang, et al.
Published: (2025)
Do-It-Yourself Automation: Interloan Bulletin Boards.
by: Moore, Cathy
Published: (1987)
by: Moore, Cathy
Published: (1987)
Do-It-Yourself Libraries
by: Dempsey, Beth
Published: (2010)
by: Dempsey, Beth
Published: (2010)
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Manifold-based Sampling for In-Context Hallucination Detection in Large Language Models
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
by: Vamshi, Bodla Krishna, et al.
Published: (2026)
MARCH: Multi-Agent Reinforced Self-Check for LLM Hallucination
by: Li, Zhuo, et al.
Published: (2026)
by: Li, Zhuo, et al.
Published: (2026)
Language Models are Bounded Pragmatic Speakers: Understanding RLHF from a Bayesian Cognitive Modeling Perspective
by: Nguyen, Khanh
Published: (2023)
by: Nguyen, Khanh
Published: (2023)
Score Before You Speak: Improving Persona Consistency in Dialogue Generation using Response Quality Scores
by: Saggar, Arpita, et al.
Published: (2025)
by: Saggar, Arpita, et al.
Published: (2025)
ComputerTown. A Do-It-Yourself Community Computer Project.
by: Loop, Liza, et al.
Published: (1982)
by: Loop, Liza, et al.
Published: (1982)
Test Yourself: What Do You Know About Reading?
by: Bloomfield, Joseph
Published: (1970)
by: Bloomfield, Joseph
Published: (1970)
Enhancing AI Safety Through the Fusion of Low Rank Adapters
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
Read Before You Think: Mitigating LLM Comprehension Failures with Step-by-Step Reading
by: Han, Feijiang, et al.
Published: (2025)
by: Han, Feijiang, et al.
Published: (2025)
Similar Items
-
Measuring Moral Inconsistencies in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024) -
Safety Training Persists Through Helpfulness Optimization in LLM Agents
by: Plaut, Benjamin
Published: (2026) -
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026) -
SaGE: Evaluating Moral Consistency in Large Language Models
by: Bonagiri, Vamshi Krishna, et al.
Published: (2024) -
Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds
by: Joishy, Anish R, et al.
Published: (2025)