Path Drift in Large Reasoning Models:How First-Person Commitments Override Safety
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Yuyi, Zhan, Runzhe, Chao, Lidia S., Tao, Ailin, Wong, Derek F. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
by: Huang, Yuyi, et al.
Published: (2025)
by: Huang, Yuyi, et al.
Published: (2025)
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
by: Xu, Haoyun, et al.
Published: (2024)
by: Xu, Haoyun, et al.
Published: (2024)
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
by: Zhan, Runzhe, et al.
Published: (2024)
by: Zhan, Runzhe, et al.
Published: (2024)
Neuron-Aware Data Selection In Instruction Tuning For Large Language Models
by: Chen, Xin, et al.
Published: (2026)
by: Chen, Xin, et al.
Published: (2026)
A Survey on LLM-Generated Text Detection: Necessity, Methods, and Future Directions
by: Wu, Junchao, et al.
Published: (2023)
by: Wu, Junchao, et al.
Published: (2023)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
by: Ma, Jingkun, et al.
Published: (2024)
by: Ma, Jingkun, et al.
Published: (2024)
Who Wrote This? The Key to Zero-Shot LLM-Generated Text Detection Is GECScore
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
DetectRL: Benchmarking LLM-Generated Text Detection in Real-World Scenarios
by: Wu, Junchao, et al.
Published: (2024)
by: Wu, Junchao, et al.
Published: (2024)
Overriding Safety protections of Open-source Models
by: Kumar, Sachin
Published: (2024)
by: Kumar, Sachin
Published: (2024)
Exposing the Cracks: Vulnerabilities of Retrieval-Augmented LLM-based Machine Translation
by: Sun, Yanming, et al.
Published: (2025)
by: Sun, Yanming, et al.
Published: (2025)
The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
by: Li, Yubo, et al.
Published: (2026)
by: Li, Yubo, et al.
Published: (2026)
ExGRPO: Learning to Reason from Experience
by: Zhan, Runzhe, et al.
Published: (2025)
by: Zhan, Runzhe, et al.
Published: (2025)
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
by: Chen, Xin, et al.
Published: (2025)
by: Chen, Xin, et al.
Published: (2025)
A Two-Stage Prediction-Aware Contrastive Learning Framework for Multi-Intent NLU
by: Chen, Guanhua, et al.
Published: (2024)
by: Chen, Guanhua, et al.
Published: (2024)
Towards an AI Musician: Synthesizing Sheet Music Problems for Musical Reasoning
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
by: He, Xiaoqi, et al.
Published: (2026)
by: He, Xiaoqi, et al.
Published: (2026)
Can ChatGPT Really Understand Modern Chinese Poetry?
by: Wang, Shanshan, et al.
Published: (2026)
by: Wang, Shanshan, et al.
Published: (2026)
What is the Best Way for ChatGPT to Translate Poetry?
by: Wang, Shanshan, et al.
Published: (2024)
by: Wang, Shanshan, et al.
Published: (2024)
Unveiling LLMs' Metaphorical Understanding: Exploring Conceptual Irrelevance, Context Leveraging and Syntactic Influence
by: Ye, Fengying, et al.
Published: (2025)
by: Ye, Fengying, et al.
Published: (2025)
FOCUS: Forging Originality through Contrastive Use in Self-Plagiarism for Language Models
by: Lan, Kaixin, et al.
Published: (2024)
by: Lan, Kaixin, et al.
Published: (2024)
How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study
by: Zhang, Zhexin, et al.
Published: (2025)
by: Zhang, Zhexin, et al.
Published: (2025)
Nevermind: Instruction Override and Moderation in Large Language Models
by: Kim, Edward
Published: (2024)
by: Kim, Edward
Published: (2024)
SGIC: A Self-Guided Iterative Calibration Framework for RAG
by: Chen, Guanhua, et al.
Published: (2025)
by: Chen, Guanhua, et al.
Published: (2025)
Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA
by: Chen, Guanhua, et al.
Published: (2026)
by: Chen, Guanhua, et al.
Published: (2026)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
by: Tam, Zhi Rui, et al.
Published: (2025)
by: Tam, Zhi Rui, et al.
Published: (2025)
Investigating CoT Monitorability in Large Reasoning Models
by: Yang, Shu, et al.
Published: (2025)
by: Yang, Shu, et al.
Published: (2025)
From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG
by: Chen, Guanhua, et al.
Published: (2026)
by: Chen, Guanhua, et al.
Published: (2026)
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning
by: Zhang, Yanfang, et al.
Published: (2024)
by: Zhang, Yanfang, et al.
Published: (2024)
Reasoning Meets Personalization: Unleashing the Potential of Large Reasoning Model for Personalized Generation
by: Luo, Sichun, et al.
Published: (2025)
by: Luo, Sichun, et al.
Published: (2025)
Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning Model
by: Lou, Xinyue, et al.
Published: (2025)
by: Lou, Xinyue, et al.
Published: (2025)
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist
by: Zhou, Zihao, et al.
Published: (2024)
by: Zhou, Zihao, et al.
Published: (2024)
Not All LoRA Parameters Are Essential: Insights on Inference Necessity
by: Chen, Guanhua, et al.
Published: (2025)
by: Chen, Guanhua, et al.
Published: (2025)
Educational Personalized Learning Path Planning with Large Language Models
by: Ng, Chee, et al.
Published: (2024)
by: Ng, Chee, et al.
Published: (2024)
Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models
by: Wang, Rui, et al.
Published: (2025)
by: Wang, Rui, et al.
Published: (2025)
Safety in Large Reasoning Models: A Survey
by: Wang, Cheng, et al.
Published: (2025)
by: Wang, Cheng, et al.
Published: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
by: Wang, Shanshan, et al.
Published: (2025)
by: Wang, Shanshan, et al.
Published: (2025)
The Better Angels of Machine Personality: How Personality Relates to LLM Safety
by: Zhang, Jie, et al.
Published: (2024)
by: Zhang, Jie, et al.
Published: (2024)
Humanity in AI: Detecting the Personality of Large Language Models
by: Zhan, Baohua, et al.
Published: (2024)
by: Zhan, Baohua, et al.
Published: (2024)
Similar Items
-
Intrinsic Model Weaknesses: How Priming Attacks Unveil Vulnerabilities in Large Language Models
by: Huang, Yuyi, et al.
Published: (2025) -
Let's Focus on Neuron: Neuron-Level Supervised Fine-tuning for Large Language Model
by: Xu, Haoyun, et al.
Published: (2024) -
Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
by: Zhan, Runzhe, et al.
Published: (2025) -
Rethinking Prompt-based Debiasing in Large Language Models
by: Yang, Xinyi, et al.
Published: (2025) -
Prefix Text as a Yarn: Eliciting Non-English Alignment in Foundation Language Model
by: Zhan, Runzhe, et al.
Published: (2024)