Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Damirchi, Hamed, De la Jara, Ignacio Meza, Abbasnejad, Ehsan, Shamsi, Afshar, Zhang, Zhen, Shi, Javen |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Decomposing Task Vectors for Refined Model Editing
par: Damirchi, Hamed, et autres
Publié: (2025)
par: Damirchi, Hamed, et autres
Publié: (2025)
The Quest for Winning Tickets in Low-Rank Adapters
par: Damirchi, Hamed, et autres
Publié: (2025)
par: Damirchi, Hamed, et autres
Publié: (2025)
Semantic Role Labeling Guided Out-of-distribution Detection
par: Zou, Jinan, et autres
Publié: (2023)
par: Zou, Jinan, et autres
Publié: (2023)
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
par: Mohammadi, Bahram, et autres
Publié: (2025)
par: Mohammadi, Bahram, et autres
Publié: (2025)
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
par: Cheng, Hongrong, et autres
Publié: (2024)
par: Cheng, Hongrong, et autres
Publié: (2024)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
par: Liu, Jiarui, et autres
Publié: (2025)
par: Liu, Jiarui, et autres
Publié: (2025)
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
par: Chrabąszcz, Maciej, et autres
Publié: (2026)
par: Chrabąszcz, Maciej, et autres
Publié: (2026)
What Makes a Representation Good for Single-Cell Perturbation Prediction?
par: Jiang, Wenkang, et autres
Publié: (2026)
par: Jiang, Wenkang, et autres
Publié: (2026)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
par: Shamsi, Afshar, et autres
Publié: (2024)
par: Shamsi, Afshar, et autres
Publié: (2024)
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
par: Wang, Keyu, et autres
Publié: (2025)
par: Wang, Keyu, et autres
Publié: (2025)
Representational and Behavioral Stability of Truth in Large Language Models
par: Dies, Samantha, et autres
Publié: (2025)
par: Dies, Samantha, et autres
Publié: (2025)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
par: Tong, Zhao, et autres
Publié: (2026)
par: Tong, Zhao, et autres
Publié: (2026)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
par: Li, Jiangnan, et autres
Publié: (2025)
par: Li, Jiangnan, et autres
Publié: (2025)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
par: Kirk, Hannah Rose, et autres
Publié: (2024)
par: Kirk, Hannah Rose, et autres
Publié: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
par: Rodriguez-Opazo, Cristian, et autres
Publié: (2024)
par: Rodriguez-Opazo, Cristian, et autres
Publié: (2024)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
par: Trinley, Katharina, et autres
Publié: (2025)
par: Trinley, Katharina, et autres
Publié: (2025)
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
par: Zheng, Haoyu, et autres
Publié: (2026)
par: Zheng, Haoyu, et autres
Publié: (2026)
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
par: Wang, Haoran, et autres
Publié: (2026)
par: Wang, Haoran, et autres
Publié: (2026)
The Geometry of Tokens in Internal Representations of Large Language Models
par: Viswanathan, Karthik, et autres
Publié: (2025)
par: Viswanathan, Karthik, et autres
Publié: (2025)
Learning Latent Dynamical Causal Processes for Single-Cell Perturbation Prediction
par: Jiang, Wenkang, et autres
Publié: (2026)
par: Jiang, Wenkang, et autres
Publié: (2026)
Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning
par: Liu, Shuo, et autres
Publié: (2026)
par: Liu, Shuo, et autres
Publié: (2026)
Probing Internal Representations of Multi-Word Verbs in Large Language Models
par: Kissane, Hassane, et autres
Publié: (2025)
par: Kissane, Hassane, et autres
Publié: (2025)
Precise Reasoning About Container-Internal Pointers with Logical Pinning
par: Guan, Yawen, et autres
Publié: (2025)
par: Guan, Yawen, et autres
Publié: (2025)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
par: Bhatia, Gagan, et autres
Publié: (2026)
par: Bhatia, Gagan, et autres
Publié: (2026)
Balancing Stylization and Truth via Disentangled Representation Steering
par: Shen, Chenglei, et autres
Publié: (2025)
par: Shen, Chenglei, et autres
Publié: (2025)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
par: Noels, Sander, et autres
Publié: (2025)
par: Noels, Sander, et autres
Publié: (2025)
When Names Disappear: Revealing What LLMs Actually Understand About Code
par: Le, Cuong Chi, et autres
Publié: (2025)
par: Le, Cuong Chi, et autres
Publié: (2025)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
par: Li, Yaoming, et autres
Publié: (2026)
par: Li, Yaoming, et autres
Publié: (2026)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
par: Yamashita, Tomoya, et autres
Publié: (2025)
par: Yamashita, Tomoya, et autres
Publié: (2025)
InvariantStock: Learning Invariant Features for Mastering the Shifting Market
par: Cao, Haiyao, et autres
Publié: (2024)
par: Cao, Haiyao, et autres
Publié: (2024)
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
par: Xu, Qiancheng, et autres
Publié: (2026)
par: Xu, Qiancheng, et autres
Publié: (2026)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
par: Cheang, Chi Seng, et autres
Publié: (2025)
par: Cheang, Chi Seng, et autres
Publié: (2025)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
par: Ghosh, Shrestha, et autres
Publié: (2025)
par: Ghosh, Shrestha, et autres
Publié: (2025)
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
par: Mousavi, Seyed Mahed, et autres
Publié: (2025)
par: Mousavi, Seyed Mahed, et autres
Publié: (2025)
The Reasoning Error About Reasoning: Why Different Types of Reasoning Require Different Representational Structures
par: Wu, Yiling
Publié: (2026)
par: Wu, Yiling
Publié: (2026)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
par: Song, Yusheng, et autres
Publié: (2025)
par: Song, Yusheng, et autres
Publié: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
par: Shamsi, Zafir, et autres
Publié: (2026)
par: Shamsi, Zafir, et autres
Publié: (2026)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
par: Winata, Genta Indra, et autres
Publié: (2026)
par: Winata, Genta Indra, et autres
Publié: (2026)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
par: Zhang, Shaolei, et autres
Publié: (2024)
par: Zhang, Shaolei, et autres
Publié: (2024)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
par: Conde, Javier, et autres
Publié: (2025)
par: Conde, Javier, et autres
Publié: (2025)
Documents similaires
-
Decomposing Task Vectors for Refined Model Editing
par: Damirchi, Hamed, et autres
Publié: (2025) -
The Quest for Winning Tickets in Low-Rank Adapters
par: Damirchi, Hamed, et autres
Publié: (2025) -
Semantic Role Labeling Guided Out-of-distribution Detection
par: Zou, Jinan, et autres
Publié: (2023) -
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
par: Mohammadi, Bahram, et autres
Publié: (2025) -
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
par: Cheng, Hongrong, et autres
Publié: (2024)