The Straight and Narrow: Do LLMs Possess an Internal Moral Path?
Fuente:
arXiv
Guardado en:
| Autores principales: | Hu, Luoming, Zeng, Jingjie, Yang, Liang, Lin, Hongfei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Large Language Models Possess Sensitive to Sentiment?
por: Liu, Yang, et al.
Publicado: (2024)
por: Liu, Yang, et al.
Publicado: (2024)
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
por: Bai, Zewen, et al.
Publicado: (2025)
por: Bai, Zewen, et al.
Publicado: (2025)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
por: Huang, Jen-tse, et al.
Publicado: (2026)
por: Huang, Jen-tse, et al.
Publicado: (2026)
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2025)
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2025)
Acquisition of Recursive Possessives and Recursive Locatives in Mandarin
por: Fu, Chenxi, et al.
Publicado: (2024)
por: Fu, Chenxi, et al.
Publicado: (2024)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
por: Zeng, Zhiyuan, et al.
Publicado: (2025)
por: Zeng, Zhiyuan, et al.
Publicado: (2025)
Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?
por: Yao, Yang, et al.
Publicado: (2025)
por: Yao, Yang, et al.
Publicado: (2025)
InternalInspector $I^2$: Robust Confidence Estimation in LLMs through Internal States
por: Beigi, Mohammad, et al.
Publicado: (2024)
por: Beigi, Mohammad, et al.
Publicado: (2024)
Commonality and Individuality! Integrating Humor Commonality with Speaker Individuality for Humor Recognition
por: Zhu, Haohao, et al.
Publicado: (2025)
por: Zhu, Haohao, et al.
Publicado: (2025)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
por: Cheang, Chi Seng, et al.
Publicado: (2025)
por: Cheang, Chi Seng, et al.
Publicado: (2025)
Common 7B Language Models Already Possess Strong Math Capabilities
por: Li, Chen, et al.
Publicado: (2024)
por: Li, Chen, et al.
Publicado: (2024)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
por: Zhou, Jingyan, et al.
Publicado: (2023)
por: Zhou, Jingyan, et al.
Publicado: (2023)
An ERP Study of Recursive Possessive Parsing in ASD Children and Its Cognitive Neuro Mechanisms
por: Chenxi, Fu, et al.
Publicado: (2026)
por: Chenxi, Fu, et al.
Publicado: (2026)
Integrating Multi-view Analysis: Multi-view Mixture-of-Expert for Textual Personality Detection
por: Zhu, Haohao, et al.
Publicado: (2024)
por: Zhu, Haohao, et al.
Publicado: (2024)
Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks
por: Bai, Zewen, et al.
Publicado: (2025)
por: Bai, Zewen, et al.
Publicado: (2025)
Internal Chain-of-Thought: Empirical Evidence for Layer-wise Subtask Scheduling in LLMs
por: Yang, Zhipeng, et al.
Publicado: (2025)
por: Yang, Zhipeng, et al.
Publicado: (2025)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
Do Large Language Models Possess a Theory of Mind? A Comparative Evaluation Using the Strange Stories Paradigm
por: Babarczy, Anna, et al.
Publicado: (2026)
por: Babarczy, Anna, et al.
Publicado: (2026)
Think Straight, Stop Smart: Structured Reasoning for Efficient Multi-Hop RAG
por: Bang, Jihwan, et al.
Publicado: (2025)
por: Bang, Jihwan, et al.
Publicado: (2025)
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
por: Wang, Haoran, et al.
Publicado: (2026)
por: Wang, Haoran, et al.
Publicado: (2026)
Children's Acquisition of Tail-recursion Sequences: A Review of Locative Recursion and Possessive Recursion as Examples
por: Wang, Xiaoyi, et al.
Publicado: (2024)
por: Wang, Xiaoyi, et al.
Publicado: (2024)
Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
por: Zhao, Siyan, et al.
Publicado: (2025)
por: Zhao, Siyan, et al.
Publicado: (2025)
Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
por: Liu, Guangliang, et al.
Publicado: (2024)
por: Liu, Guangliang, et al.
Publicado: (2024)
Straight to Zero: Why Linearly Decaying the Learning Rate to Zero Works Best for LLMs
por: Bergsma, Shane, et al.
Publicado: (2025)
por: Bergsma, Shane, et al.
Publicado: (2025)
Probing the Lack of Stable Internal Beliefs in LLMs
por: Luo, Yifan, et al.
Publicado: (2026)
por: Luo, Yifan, et al.
Publicado: (2026)
Is LLM an Overconfident Judge? Unveiling the Capabilities of LLMs in Detecting Offensive Language with Annotation Disagreement
por: Lu, Junyu, et al.
Publicado: (2025)
por: Lu, Junyu, et al.
Publicado: (2025)
Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
por: Xiao, Kelaiti, et al.
Publicado: (2025)
por: Xiao, Kelaiti, et al.
Publicado: (2025)
Moral Mazes in the Era of LLMs
por: Nguyen, Dang, et al.
Publicado: (2026)
por: Nguyen, Dang, et al.
Publicado: (2026)
AI Research Agents Narrow Scientific Exploration
por: Tang, Yixuan, et al.
Publicado: (2026)
por: Tang, Yixuan, et al.
Publicado: (2026)
IPEval: A Bilingual Intellectual Property Agency Consultation Evaluation Benchmark for Large Language Models
por: Wang, Qiyao, et al.
Publicado: (2024)
por: Wang, Qiyao, et al.
Publicado: (2024)
HearSay Benchmark: Do Audio LLMs Leak What They Hear?
por: Wang, Jin, et al.
Publicado: (2026)
por: Wang, Jin, et al.
Publicado: (2026)
How Well Do LLMs Identify Cultural Unity in Diversity?
por: Li, Jialin, et al.
Publicado: (2024)
por: Li, Jialin, et al.
Publicado: (2024)
APAR: LLMs Can Do Auto-Parallel Auto-Regressive Decoding
por: Liu, Mingdao, et al.
Publicado: (2024)
por: Liu, Mingdao, et al.
Publicado: (2024)
RexDrug: Reliable Multi-Drug Combination Extraction through Reasoning-Enhanced LLMs
por: Wang, Zhijun, et al.
Publicado: (2026)
por: Wang, Zhijun, et al.
Publicado: (2026)
Do Emotions Influence Moral Judgment in Large Language Models?
por: Saim, Mohammad, et al.
Publicado: (2026)
por: Saim, Mohammad, et al.
Publicado: (2026)
Do Large Language Models Understand Morality Across Cultures?
por: Mohammadi, Hadi, et al.
Publicado: (2025)
por: Mohammadi, Hadi, et al.
Publicado: (2025)
Targeted Distillation for Sentiment Analysis
por: Zhang, Yice, et al.
Publicado: (2025)
por: Zhang, Yice, et al.
Publicado: (2025)
Hallucination Detection with the Internal Layers of LLMs
por: Preiß, Martin
Publicado: (2025)
por: Preiß, Martin
Publicado: (2025)
Seeing Straight: Document Orientation Detection for Efficient OCR
por: Goswami, Suranjan, et al.
Publicado: (2025)
por: Goswami, Suranjan, et al.
Publicado: (2025)
Evaluating Gender Bias of LLMs in Making Morality Judgements
por: Bajaj, Divij, et al.
Publicado: (2024)
por: Bajaj, Divij, et al.
Publicado: (2024)
Ejemplares similares
-
Do Large Language Models Possess Sensitive to Sentiment?
por: Liu, Yang, et al.
Publicado: (2024) -
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
por: Bai, Zewen, et al.
Publicado: (2025) -
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
por: Huang, Jen-tse, et al.
Publicado: (2026) -
Do Internal Layers of LLMs Reveal Patterns for Jailbreak Detection?
por: Kadali, Sri Durga Sai Sowmya, et al.
Publicado: (2025) -
Acquisition of Recursive Possessives and Recursive Locatives in Mandarin
por: Fu, Chenxi, et al.
Publicado: (2024)