Easier to Mislead Than to Correct: Harmful and Beneficial Revision in LLM Conformity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qu, Jiaming, fu, Lucheng, Hu, Yibo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
von: Dang, Yunkai, et al.
Veröffentlicht: (2024)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
von: Yang, Langqi, et al.
Veröffentlicht: (2025)
Self-HarmLLM: Can Large Language Model Harm Itself?
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
von: Kim, Heehwan, et al.
Veröffentlicht: (2025)
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
von: Chen, Zixin, et al.
Veröffentlicht: (2025)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
von: Lo, Leo Yu-Ho, et al.
Veröffentlicht: (2024)
von: Lo, Leo Yu-Ho, et al.
Veröffentlicht: (2024)
Mitigating Misleading Chain-of-Thought Reasoning with Selective Filtering
von: Wu, Yexin, et al.
Veröffentlicht: (2024)
von: Wu, Yexin, et al.
Veröffentlicht: (2024)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Large Language Models as Misleading Assistants in Conversation
von: Hou, Betty Li, et al.
Veröffentlicht: (2024)
von: Hou, Betty Li, et al.
Veröffentlicht: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
von: Pandey, Punya Syon, et al.
Veröffentlicht: (2025)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2026)
von: Sun, Zhongxiang, et al.
Veröffentlicht: (2026)
HALF: Harm-Aware LLM Fairness Evaluation Aligned with Deployment
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
von: Mekky, Ali, et al.
Veröffentlicht: (2025)
Booster: Tackling Harmful Fine-tuning for Large Language Models via Attenuating Harmful Perturbation
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
von: Huang, Tiansheng, et al.
Veröffentlicht: (2024)
Metaphors We Compute By: A Computational Audit of Cultural Translation vs. Thinking in LLMs
von: Chang, Yuan, et al.
Veröffentlicht: (2026)
von: Chang, Yuan, et al.
Veröffentlicht: (2026)
REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
von: Jindal, Madhur, et al.
Veröffentlicht: (2025)
von: Jindal, Madhur, et al.
Veröffentlicht: (2025)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Learning Shortcuts: On the Misleading Promise of NLU in Language Models
von: Bihani, Geetanjali, et al.
Veröffentlicht: (2024)
von: Bihani, Geetanjali, et al.
Veröffentlicht: (2024)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
von: Sharshar, Ahmed, et al.
Veröffentlicht: (2026)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Trivial Vocabulary Bans Improve LLM Reasoning More Than Deep Linguistic Constraints
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
von: Jehu-Appiah, Rodney
Veröffentlicht: (2026)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
von: Zhang, Shaoqing, et al.
Veröffentlicht: (2024)
von: Zhang, Shaoqing, et al.
Veröffentlicht: (2024)
HarmTransform: Transforming Explicit Harmful Queries into Stealthy via Multi-Agent Debate
von: Zhu, Shenzhe
Veröffentlicht: (2025)
von: Zhu, Shenzhe
Veröffentlicht: (2025)
More Than Sum of Its Parts: Deciphering Intent Shifts in Multimodal Hate Speech Detection
von: Sun, Runze, et al.
Veröffentlicht: (2026)
von: Sun, Runze, et al.
Veröffentlicht: (2026)
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring
von: Chen, Guanxu, et al.
Veröffentlicht: (2025)
von: Chen, Guanxu, et al.
Veröffentlicht: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
`For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
von: Schoene, Annika M, et al.
Veröffentlicht: (2025)
An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
von: Kaur, Navdeep, et al.
Veröffentlicht: (2025)
Enhancing LLM-based Hatred and Toxicity Detection with Meta-Toxic Knowledge Graph
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
von: Zhao, Yibo, et al.
Veröffentlicht: (2024)
KLong: Training LLM Agent for Extremely Long-horizon Tasks
von: Liu, Yue, et al.
Veröffentlicht: (2026)
von: Liu, Yue, et al.
Veröffentlicht: (2026)
Towards Comprehensive Detection of Chinese Harmful Memes
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
von: Lu, Junyu, et al.
Veröffentlicht: (2024)
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
von: Kim, Jinhwa, et al.
Veröffentlicht: (2025)
von: Kim, Jinhwa, et al.
Veröffentlicht: (2025)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
von: Liu, Zijie, et al.
Veröffentlicht: (2025)
Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning
von: Xia, Haizhou
Veröffentlicht: (2026)
von: Xia, Haizhou
Veröffentlicht: (2026)
VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
von: Sun, Qixin, et al.
Veröffentlicht: (2025)
von: Sun, Qixin, et al.
Veröffentlicht: (2025)
CHARTOM: A Visual Theory-of-Mind Benchmark for LLMs on Misleading Charts
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
von: Bharti, Shubham, et al.
Veröffentlicht: (2024)
Human-Guided Harm Recovery for Computer Use Agents
von: Li, Christy, et al.
Veröffentlicht: (2026)
von: Li, Christy, et al.
Veröffentlicht: (2026)
FairBelief -- Assessing Harmful Beliefs in Language Models
von: Setzu, Mattia, et al.
Veröffentlicht: (2024)
von: Setzu, Mattia, et al.
Veröffentlicht: (2024)
Mitigating LLM Hallucinations via Conformal Abstention
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
von: Yadkori, Yasin Abbasi, et al.
Veröffentlicht: (2024)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2024)
Two Minds Better Than One: Collaborative Reward Modeling for LLM Alignment
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jiazheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
von: Dang, Yunkai, et al.
Veröffentlicht: (2024) -
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
von: Yang, Langqi, et al.
Veröffentlicht: (2025) -
Self-HarmLLM: Can Large Language Model Harm Itself?
von: Kim, Heehwan, et al.
Veröffentlicht: (2025) -
Unmasking Deceptive Visuals: Benchmarking Multimodal Large Language Models on Misleading Chart Question Answering
von: Chen, Zixin, et al.
Veröffentlicht: (2025) -
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
von: Lo, Leo Yu-Ho, et al.
Veröffentlicht: (2024)