Understanding the Dark Side of LLMs' Intrinsic Self-Correction
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Qingjie, Wang, Di, Qian, Haoting, Li, Yiming, Zhang, Tianwei, Huang, Minlie, Xu, Ke, Li, Hewu, Liu, Yan, Qiu, Han |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
Understanding the Dilemma of Unlearning for Large Language Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
di: Dong, Jianshuo, et al.
Pubblicazione: (2025)
di: Dong, Jianshuo, et al.
Pubblicazione: (2025)
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
di: Liu, Guangliang, et al.
Pubblicazione: (2024)
di: Liu, Guangliang, et al.
Pubblicazione: (2024)
An Engorgio Prompt Makes Large Language Model Babble on
di: Dong, Jianshuo, et al.
Pubblicazione: (2024)
di: Dong, Jianshuo, et al.
Pubblicazione: (2024)
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
di: Kamoi, Ryo, et al.
Pubblicazione: (2024)
di: Kamoi, Ryo, et al.
Pubblicazione: (2024)
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
di: Yang, Shujian, et al.
Pubblicazione: (2025)
di: Yang, Shujian, et al.
Pubblicazione: (2025)
ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
di: Zhang, Zhexin, et al.
Pubblicazione: (2024)
Course-Correction: Safety Alignment Using Synthetic Preferences
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
di: Chen, Renmiao, et al.
Pubblicazione: (2026)
Self-Correction Makes LLMs Better Parsers
di: Zhang, Ziyan, et al.
Pubblicazione: (2025)
di: Zhang, Ziyan, et al.
Pubblicazione: (2025)
Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
di: Li, Loka, et al.
Pubblicazione: (2024)
di: Li, Loka, et al.
Pubblicazione: (2024)
When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' Toxicity
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
di: Cui, Shiyao, et al.
Pubblicazione: (2025)
Large Language Models have Intrinsic Self-Correction Ability
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
di: Liu, Dancheng, et al.
Pubblicazione: (2024)
DSB: Dynamic Sliding Block Scheduling for Diffusion LLMs
di: Luo, Lizhuo, et al.
Pubblicazione: (2026)
di: Luo, Lizhuo, et al.
Pubblicazione: (2026)
Picky LLMs and Unreliable RMs: An Empirical Study on Safety Alignment after Instruction Tuning
di: Li, Guanlin, et al.
Pubblicazione: (2025)
di: Li, Guanlin, et al.
Pubblicazione: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
di: Yang, Xin, et al.
Pubblicazione: (2026)
di: Yang, Xin, et al.
Pubblicazione: (2026)
Silenced Biases: The Dark Side LLMs Learned to Refuse
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
A Case for Application-Aware Space Radiation Tolerance in Orbital Computing
di: Wang, Meiqi, et al.
Pubblicazione: (2024)
di: Wang, Meiqi, et al.
Pubblicazione: (2024)
Seeker: Towards Exception Safety Code Generation with Intermediate Language Agents Framework
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
di: Lee, Yu-Ting, et al.
Pubblicazione: (2025)
Language Model Decoding as Direct Metrics Optimization
di: Ji, Haozhe, et al.
Pubblicazione: (2023)
di: Ji, Haozhe, et al.
Pubblicazione: (2023)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
di: Tie, Guiyao, et al.
Pubblicazione: (2025)
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
di: Zhang, Chenkai, et al.
Pubblicazione: (2025)
di: Zhang, Chenkai, et al.
Pubblicazione: (2025)
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
di: Liu, Songyang, et al.
Pubblicazione: (2025)
di: Liu, Songyang, et al.
Pubblicazione: (2025)
QueueEDIT: Structural Self-Correction for Sequential Model Editing in LLMs
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
COSMIC: Compress Satellite Images Efficiently via Diffusion Compensation
di: Zhang, Ziyuan, et al.
Pubblicazione: (2024)
di: Zhang, Ziyuan, et al.
Pubblicazione: (2024)
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
di: Zhang, Zhexin, et al.
Pubblicazione: (2025)
di: Zhang, Zhexin, et al.
Pubblicazione: (2025)
The Shadow Self: Intrinsic Value Misalignment in Large Language Model Agents
di: Chen, Chen, et al.
Pubblicazione: (2026)
di: Chen, Chen, et al.
Pubblicazione: (2026)
Confidence v.s. Critique: A Decomposition of Self-Correction Capability for LLMs
di: Yang, Zhe, et al.
Pubblicazione: (2024)
di: Yang, Zhe, et al.
Pubblicazione: (2024)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
Defending Large Language Models Against Jailbreaking Attacks Through Goal Prioritization
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
di: Zhang, Zhexin, et al.
Pubblicazione: (2023)
Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and Bias
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
di: Jiang, Botian, et al.
Pubblicazione: (2024)
di: Jiang, Botian, et al.
Pubblicazione: (2024)
SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding
di: Li, Sihang, et al.
Pubblicazione: (2024)
di: Li, Sihang, et al.
Pubblicazione: (2024)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
di: Zhang, Junjie, et al.
Pubblicazione: (2025)
di: Zhang, Junjie, et al.
Pubblicazione: (2025)
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
di: Zhang, Xuan, et al.
Pubblicazione: (2025)
di: Zhang, Xuan, et al.
Pubblicazione: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
di: Huang, Yue, et al.
Pubblicazione: (2025)
di: Huang, Yue, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Speculating LLMs' Chinese Training Data Pollution from Their Tokens
di: Zhang, Qingjie, et al.
Pubblicazione: (2025) -
Understanding the Dilemma of Unlearning for Large Language Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025) -
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025) -
SafeSearch: Automated Red-Teaming of LLM-Based Search Agents
di: Dong, Jianshuo, et al.
Pubblicazione: (2025) -
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
di: Liu, Guangliang, et al.
Pubblicazione: (2024)