Understanding Moral Reasoning Trajectories in Large Language Models: Toward Probing-Based Explainability
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Fan, Kwak, Haewoon, An, Jisun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
by: Huang, Fan, et al.
Published: (2024)
by: Huang, Fan, et al.
Published: (2024)
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025)
by: Muralidharan, Rasika, et al.
Published: (2025)
CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models
by: Huang, Fan, et al.
Published: (2026)
by: Huang, Fan, et al.
Published: (2026)
Can we trust the evaluation on ChatGPT?
by: Aiyappa, Rachith, et al.
Published: (2023)
by: Aiyappa, Rachith, et al.
Published: (2023)
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
by: Aiyappa, Rachith, et al.
Published: (2024)
by: Aiyappa, Rachith, et al.
Published: (2024)
XChoice: Explainable Evaluation of AI-Human Alignment in LLM-based Constrained Choice Decision Making
by: Qi, Weihong, et al.
Published: (2026)
by: Qi, Weihong, et al.
Published: (2026)
A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies
by: Qi, Weihong, et al.
Published: (2025)
by: Qi, Weihong, et al.
Published: (2025)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
by: Kachwala, Zoher, et al.
Published: (2026)
by: Kachwala, Zoher, et al.
Published: (2026)
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
Towards Explainable Temporal Reasoning in Large Language Models: A Structure-Aware Generative Framework
by: Jiang, Zihao, et al.
Published: (2025)
by: Jiang, Zihao, et al.
Published: (2025)
Rematch: Robust and Efficient Matching of Local Knowledge Graphs to Improve Structural and Semantic Similarity
by: Kachwala, Zoher, et al.
Published: (2024)
by: Kachwala, Zoher, et al.
Published: (2024)
From Understanding to Utilization: A Survey on Explainability for Large Language Models
by: Luo, Haoyan, et al.
Published: (2024)
by: Luo, Haoyan, et al.
Published: (2024)
Probing the "Psyche'' of Large Reasoning Models: Understanding Through a Human Lens
by: Chen, Yuxiang, et al.
Published: (2025)
by: Chen, Yuxiang, et al.
Published: (2025)
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
by: Nunes, José Luiz, et al.
Published: (2024)
by: Nunes, José Luiz, et al.
Published: (2024)
Stuck in the Matrix: Probing Spatial Reasoning in Large Language Models
by: Bai, Maggie, et al.
Published: (2025)
by: Bai, Maggie, et al.
Published: (2025)
Do Language Models Understand Morality? Towards a Robust Detection of Moral Content
by: Bulla, Luana, et al.
Published: (2024)
by: Bulla, Luana, et al.
Published: (2024)
AdamMeme: Adaptively Probe the Reasoning Capacity of Multimodal Large Language Models on Harmfulness
by: Chen, Zixin, et al.
Published: (2025)
by: Chen, Zixin, et al.
Published: (2025)
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
by: Xu, Fengli, et al.
Published: (2025)
by: Xu, Fengli, et al.
Published: (2025)
GraphInstruct: Empowering Large Language Models with Graph Understanding and Reasoning Capability
by: Luo, Zihan, et al.
Published: (2024)
by: Luo, Zihan, et al.
Published: (2024)
Towards Explainable Evolution Strategies with Large Language Models
by: Baumann, Jill, et al.
Published: (2024)
by: Baumann, Jill, et al.
Published: (2024)
Probabilistic Aggregation and Targeted Embedding Optimization for Collective Moral Reasoning in Large Language Models
by: Yuan, Chenchen, et al.
Published: (2025)
by: Yuan, Chenchen, et al.
Published: (2025)
What Helps Language Models Predict Human Beliefs: Demographics or Prior Stances?
by: Malone, Joseph, et al.
Published: (2025)
by: Malone, Joseph, et al.
Published: (2025)
Enhancing Large Language Models through Structured Reasoning
by: Dong, Yubo, et al.
Published: (2025)
by: Dong, Yubo, et al.
Published: (2025)
The Pluralistic Moral Gap: Understanding Judgment and Value Differences between Humans and Large Language Models
by: Russo, Giuseppe, et al.
Published: (2025)
by: Russo, Giuseppe, et al.
Published: (2025)
Moral Persuasion in Large Language Models: Evaluating Susceptibility and Ethical Alignment
by: Huang, Allison, et al.
Published: (2024)
by: Huang, Allison, et al.
Published: (2024)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
by: Ball, Sarah, et al.
Published: (2025)
by: Ball, Sarah, et al.
Published: (2025)
Towards Understanding the Cognitive Habits of Large Reasoning Models
by: Dong, Jianshuo, et al.
Published: (2025)
by: Dong, Jianshuo, et al.
Published: (2025)
Mixture of Reasonings: Teach Large Language Models to Reason with Adaptive Strategies
by: Xiong, Tao, et al.
Published: (2025)
by: Xiong, Tao, et al.
Published: (2025)
Neural Probe-Based Hallucination Detection for Large Language Models
by: Liang, Shize, et al.
Published: (2025)
by: Liang, Shize, et al.
Published: (2025)
Tracing Moral Foundations in Large Language Models
by: Yu, Chenxiao, et al.
Published: (2026)
by: Yu, Chenxiao, et al.
Published: (2026)
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability
by: Xu, Chenhui, et al.
Published: (2025)
by: Xu, Chenhui, et al.
Published: (2025)
LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?
by: Wang, Jingyuan, et al.
Published: (2025)
by: Wang, Jingyuan, et al.
Published: (2025)
Understanding the Dilemma of Unlearning for Large Language Models
by: Zhang, Qingjie, et al.
Published: (2025)
by: Zhang, Qingjie, et al.
Published: (2025)
XRec: Large Language Models for Explainable Recommendation
by: Ma, Qiyao, et al.
Published: (2024)
by: Ma, Qiyao, et al.
Published: (2024)
ReProbe: Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
by: Ni, Jingwei, et al.
Published: (2025)
by: Ni, Jingwei, et al.
Published: (2025)
BELL: Benchmarking the Explainability of Large Language Models
by: Ahmed, Syed Quiser, et al.
Published: (2025)
by: Ahmed, Syed Quiser, et al.
Published: (2025)
Towards Explainable Harmful Meme Detection through Multimodal Debate between Large Language Models
by: Lin, Hongzhan, et al.
Published: (2024)
by: Lin, Hongzhan, et al.
Published: (2024)
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding
by: Zhang, Yazhou, et al.
Published: (2024)
by: Zhang, Yazhou, et al.
Published: (2024)
Similar Items
-
Vulnerability of LLMs' Stated Beliefs? LLMs Belief Resistance Check Through Strategic Persuasive Conversation Interventions
by: Huang, Fan, et al.
Published: (2026) -
ToBlend: Token-Level Blending With an Ensemble of LLMs to Attack AI-Generated Text Detection
by: Huang, Fan, et al.
Published: (2024) -
ChatGPT Rates Natural Language Explanation Quality Like Humans: But on Which Scales?
by: Huang, Fan, et al.
Published: (2024) -
Can Lessons From Human Teams Be Applied to Multi-Agent Systems? The Role of Structure, Diversity, and Interaction Dynamics
by: Muralidharan, Rasika, et al.
Published: (2025) -
CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models
by: Huang, Fan, et al.
Published: (2026)