Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Phan, Hoang, Li, Victor, Lei, Qi |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
par: Tang, Fei, et autres
Publié: (2025)
par: Tang, Fei, et autres
Publié: (2025)
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
par: Lacombe, Romain, et autres
Publié: (2025)
par: Lacombe, Romain, et autres
Publié: (2025)
Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models
par: Mikula, Lukáš, et autres
Publié: (2023)
par: Mikula, Lukáš, et autres
Publié: (2023)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
par: He, Jiashu, et autres
Publié: (2026)
par: He, Jiashu, et autres
Publié: (2026)
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
par: Wang, Ziyan, et autres
Publié: (2025)
par: Wang, Ziyan, et autres
Publié: (2025)
Think Twice, Act Once: Verifier-Guided Action Selection For Embodied Agents
par: Singhi, Nishad, et autres
Publié: (2026)
par: Singhi, Nishad, et autres
Publié: (2026)
Tamper-Resistant Safeguards for Open-Weight LLMs
par: Tamirisa, Rishub, et autres
Publié: (2024)
par: Tamirisa, Rishub, et autres
Publié: (2024)
MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking
par: Ji, Shihao, et autres
Publié: (2025)
par: Ji, Shihao, et autres
Publié: (2025)
A Framework for Real-time Safeguarding the Text Generation of Large Language Model
par: Dong, Ximing, et autres
Publié: (2024)
par: Dong, Ximing, et autres
Publié: (2024)
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
par: Qin, Kai, et autres
Publié: (2026)
par: Qin, Kai, et autres
Publié: (2026)
PRIMO: Progressive Induction for Multi-hop Open Rule Generation
par: Liu, Jianyu, et autres
Publié: (2024)
par: Liu, Jianyu, et autres
Publié: (2024)
Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
par: Du, Chengyu, et autres
Publié: (2024)
par: Du, Chengyu, et autres
Publié: (2024)
MixReasoning: Switching Modes to Think
par: Lu, Haiquan, et autres
Publié: (2025)
par: Lu, Haiquan, et autres
Publié: (2025)
STACK: Adversarial Attacks on LLM Safeguard Pipelines
par: McKenzie, Ian R., et autres
Publié: (2025)
par: McKenzie, Ian R., et autres
Publié: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
par: Huang, Hui, et autres
Publié: (2025)
par: Huang, Hui, et autres
Publié: (2025)
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
par: Si, Shengyun, et autres
Publié: (2025)
par: Si, Shengyun, et autres
Publié: (2025)
Output Length Effect on DeepSeek-R1's Safety in Forced Thinking
par: Li, Xuying, et autres
Publié: (2025)
par: Li, Xuying, et autres
Publié: (2025)
Finetune Once: Decoupling General & Domain Learning with Dynamic Boosted Annealing
par: Tang, Yang, et autres
Publié: (2025)
par: Tang, Yang, et autres
Publié: (2025)
VNJPTranslate: A comprehensive pipeline for Vietnamese-Japanese translation
par: Phan, Hoang Hai, et autres
Publié: (2025)
par: Phan, Hoang Hai, et autres
Publié: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
par: Huang, Yue, et autres
Publié: (2025)
par: Huang, Yue, et autres
Publié: (2025)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
par: Guo, Yipin, et autres
Publié: (2024)
par: Guo, Yipin, et autres
Publié: (2024)
ThinkTuning: Instilling Cognitive Reflections without Distillation
par: RRV, Aswin, et autres
Publié: (2025)
par: RRV, Aswin, et autres
Publié: (2025)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
par: Sahoo, Subramanyam, et autres
Publié: (2026)
par: Sahoo, Subramanyam, et autres
Publié: (2026)
Calibrating Verbalized Confidence with Self-Generated Distractors
par: Wang, Victor, et autres
Publié: (2025)
par: Wang, Victor, et autres
Publié: (2025)
MCTSr-Zero: Self-Reflective Psychological Counseling Dialogues Generation via Principles and Adaptive Exploration
par: Lu, Hao, et autres
Publié: (2025)
par: Lu, Hao, et autres
Publié: (2025)
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation
par: Ren, Houxing, et autres
Publié: (2024)
par: Ren, Houxing, et autres
Publié: (2024)
Thinking Out of Order: When Output Order Stops Reflecting Reasoning Order in Diffusion Language Models
par: Yu, Longxuan, et autres
Publié: (2026)
par: Yu, Longxuan, et autres
Publié: (2026)
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
par: Pang, Renning, et autres
Publié: (2026)
par: Pang, Renning, et autres
Publié: (2026)
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning
par: Yu, Simon, et autres
Publié: (2024)
par: Yu, Simon, et autres
Publié: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
par: Zhang, Jiajie, et autres
Publié: (2025)
par: Zhang, Jiajie, et autres
Publié: (2025)
Towards Trustworthy Report Generation: A Deep Research Agent with Progressive Confidence Estimation and Calibration
par: Yuan, Yi, et autres
Publié: (2026)
par: Yuan, Yi, et autres
Publié: (2026)
Thinking LLMs: General Instruction Following with Thought Generation
par: Wu, Tianhao, et autres
Publié: (2024)
par: Wu, Tianhao, et autres
Publié: (2024)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
par: Zhang, Dalong, et autres
Publié: (2025)
par: Zhang, Dalong, et autres
Publié: (2025)
Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
par: Wang, Jian, et autres
Publié: (2024)
par: Wang, Jian, et autres
Publié: (2024)
sDPO: Don't Use Your Data All at Once
par: Kim, Dahyun, et autres
Publié: (2024)
par: Kim, Dahyun, et autres
Publié: (2024)
Edit Once, Update Everywhere: A Simple Framework for Cross-Lingual Knowledge Synchronization in LLMs
par: Wu, Yuchen, et autres
Publié: (2025)
par: Wu, Yuchen, et autres
Publié: (2025)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
par: Li, Moxin, et autres
Publié: (2024)
par: Li, Moxin, et autres
Publié: (2024)
Reinforcement Learning from Reflective Feedback (RLRF): Aligning and Improving LLMs via Fine-Grained Self-Reflection
par: Lee, Kyungjae, et autres
Publié: (2024)
par: Lee, Kyungjae, et autres
Publié: (2024)
When to Trust Context: Self-Reflective Debates for Context Reliability
par: Zhou, Zeqi, et autres
Publié: (2025)
par: Zhou, Zeqi, et autres
Publié: (2025)
Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
par: Zhang, Wenqi, et autres
Publié: (2024)
par: Zhang, Wenqi, et autres
Publié: (2024)
Documents similaires
-
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
par: Tang, Fei, et autres
Publié: (2025) -
Don't Think Twice! Over-Reasoning Impairs Confidence Calibration
par: Lacombe, Romain, et autres
Publié: (2025) -
Think Twice: Measuring the Efficiency of Eliminating Prediction Shortcuts of Question Answering Models
par: Mikula, Lukáš, et autres
Publié: (2023) -
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
par: He, Jiashu, et autres
Publié: (2026) -
Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
par: Wang, Ziyan, et autres
Publié: (2025)