Strengthening Human-Centric Chain-of-Thought Reasoning Integrity in LLMs via a Structured Prompt Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Jiling, Adeseye, Aisvarya, Virtanen, Seppo, Hakkala, Antti, Isoaho, Jouni |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Local Language Models for Context-Aware Adaptive Anonymization of Sensitive Text
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
Dynamic Agentic AI Expert Profiler System Architecture for Multidomain Intelligence Modeling
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
Modular AI-Powered Interviewer with Dynamic Question Generation and Expertise Profiling
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2025)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2025)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)
Entropy and Attention Dynamics in Small Language Models: A Trace-Level Structural Analysis on the TruthfulQA Benchmark
di: Adeseye, Adeyemi, et al.
Pubblicazione: (2026)
di: Adeseye, Adeyemi, et al.
Pubblicazione: (2026)
Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning
di: Chang, Zhiyuan, et al.
Pubblicazione: (2026)
di: Chang, Zhiyuan, et al.
Pubblicazione: (2026)
Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
di: Zolkowski, Artur, et al.
Pubblicazione: (2025)
Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
di: Lu, Yu-An, et al.
Pubblicazione: (2026)
di: Lu, Yu-An, et al.
Pubblicazione: (2026)
A Prompt-Based Framework for Loop Vulnerability Detection Using Local LLMs
di: Adeseye, Adeyemi, et al.
Pubblicazione: (2026)
di: Adeseye, Adeyemi, et al.
Pubblicazione: (2026)
LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
di: Li, Chloe, et al.
Pubblicazione: (2025)
di: Li, Chloe, et al.
Pubblicazione: (2025)
Preemptive Answer "Attacks" on Chain-of-Thought Reasoning
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
di: Xu, Rongwu, et al.
Pubblicazione: (2024)
BadThink: Triggered Overthinking Attacks on Chain-of-Thought Reasoning in Large Language Models
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
di: Liu, Shuaitong, et al.
Pubblicazione: (2025)
Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
di: MacDermott, Matt, et al.
Pubblicazione: (2025)
di: MacDermott, Matt, et al.
Pubblicazione: (2025)
ExplainableGuard: Interpretable Adversarial Defense for Large Language Models Using Chain-of-Thought Reasoning
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
di: Guan, Shaowei, et al.
Pubblicazione: (2025)
SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration
di: Pan, Yu, et al.
Pubblicazione: (2026)
di: Pan, Yu, et al.
Pubblicazione: (2026)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
di: Bae, Yubeen, et al.
Pubblicazione: (2025)
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture
di: Xiang, Rong
Pubblicazione: (2026)
di: Xiang, Rong
Pubblicazione: (2026)
Large Language Model-driven Security Assistant for Internet of Things via Chain-of-Thought
di: Zeng, Mingfei, et al.
Pubblicazione: (2025)
di: Zeng, Mingfei, et al.
Pubblicazione: (2025)
Enhancing Jailbreak Attacks on LLMs via Persona Prompts
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
di: Zhang, Zheng, et al.
Pubblicazione: (2025)
PromptKeeper: Safeguarding System Prompts for LLMs
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
di: Jiang, Zhifeng, et al.
Pubblicazione: (2024)
From Thinking to Output: Chain-of-Thought and Text Generation Characteristics in Reasoning Language Models
di: Liu, Junhao, et al.
Pubblicazione: (2025)
di: Liu, Junhao, et al.
Pubblicazione: (2025)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
di: Mireshghallah, Niloofar, et al.
Pubblicazione: (2023)
Prompt Flow Integrity to Prevent Privilege Escalation in LLM Agents
di: Kim, Juhee, et al.
Pubblicazione: (2025)
di: Kim, Juhee, et al.
Pubblicazione: (2025)
Hey GPT-OSS, Looks Like You Got It -- Now Walk Me Through It! An Assessment of the Reasoning Language Models Chain of Thought Mechanism for Digital Forensics
di: Michelet, Gaëtan, et al.
Pubblicazione: (2025)
di: Michelet, Gaëtan, et al.
Pubblicazione: (2025)
CoTSRF: Utilize Chain of Thought as Stealthy and Robust Fingerprint of Large Language Models
di: Ren, Zhenzhen, et al.
Pubblicazione: (2025)
di: Ren, Zhenzhen, et al.
Pubblicazione: (2025)
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
di: Roh, Jaechul, et al.
Pubblicazione: (2025)
Decoding Latent Attack Surfaces in LLMs: Prompt Injection via HTML in Web Summarization
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
di: Verma, Ishaan, et al.
Pubblicazione: (2025)
Chain-of-Thought Prompting of Large Language Models for Discovering and Fixing Software Vulnerabilities
di: Nong, Yu, et al.
Pubblicazione: (2024)
di: Nong, Yu, et al.
Pubblicazione: (2024)
Output Supervision Can Obfuscate the Chain of Thought
di: Drori, Jacob, et al.
Pubblicazione: (2025)
di: Drori, Jacob, et al.
Pubblicazione: (2025)
Imitation Game for Adversarial Disillusion with Chain-of-Thought Reasoning in Generative AI
di: Chang, Ching-Chun, et al.
Pubblicazione: (2025)
di: Chang, Ching-Chun, et al.
Pubblicazione: (2025)
Unreal Thinking: Chain-of-Thought Hijacking via Two-stage Backdoor
di: Chang, Wenhan, et al.
Pubblicazione: (2026)
di: Chang, Wenhan, et al.
Pubblicazione: (2026)
It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMs
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
di: Park, Sangwoo, et al.
Pubblicazione: (2026)
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
di: Xiang, Zhen, et al.
Pubblicazione: (2024)
Thought Purity: A Defense Framework For Chain-of-Thought Attack
di: Xue, Zihao, et al.
Pubblicazione: (2025)
di: Xue, Zihao, et al.
Pubblicazione: (2025)
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
di: Jaiswal, Piyush, et al.
Pubblicazione: (2026)
Prompt and Circumstances: Evaluating the Efficacy of Human Prompt Inference in AI-Generated Art
di: Trinh, Khoi, et al.
Pubblicazione: (2026)
di: Trinh, Khoi, et al.
Pubblicazione: (2026)
Taxonomy, Evaluation and Exploitation of IPI-Centric LLM Agent Defense Frameworks
di: Ji, Zimo, et al.
Pubblicazione: (2025)
di: Ji, Zimo, et al.
Pubblicazione: (2025)
R-CoT: A Reasoning-Layer Watermark via Redundant Chain-of-Thought in Large Language Models
di: Zhang, Ziming, et al.
Pubblicazione: (2026)
di: Zhang, Ziming, et al.
Pubblicazione: (2026)
Rethinking Reasoning: A Survey on Reasoning-based Backdoors in LLMs
di: Hu, Man, et al.
Pubblicazione: (2025)
di: Hu, Man, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Local Language Models for Context-Aware Adaptive Anonymization of Sensitive Text
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026) -
Dynamic Agentic AI Expert Profiler System Architecture for Multidomain Intelligence Modeling
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026) -
Modular AI-Powered Interviewer with Dynamic Question Generation and Expertise Profiling
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2025) -
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026) -
Parallel LLM Reasoning for Bias-Resilient, Robust Conceptual Abstraction
di: Adeseye, Aisvarya, et al.
Pubblicazione: (2026)