Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
Fuente:
arXiv
Guardado en:
| Autores principales: | Damirchi, Hamed, De la Jara, Ignacio Meza, Abbasnejad, Ehsan, Shamsi, Afshar, Zhang, Zhen, Shi, Javen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Decomposing Task Vectors for Refined Model Editing
por: Damirchi, Hamed, et al.
Publicado: (2025)
por: Damirchi, Hamed, et al.
Publicado: (2025)
The Quest for Winning Tickets in Low-Rank Adapters
por: Damirchi, Hamed, et al.
Publicado: (2025)
por: Damirchi, Hamed, et al.
Publicado: (2025)
Semantic Role Labeling Guided Out-of-distribution Detection
por: Zou, Jinan, et al.
Publicado: (2023)
por: Zou, Jinan, et al.
Publicado: (2023)
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
por: Mohammadi, Bahram, et al.
Publicado: (2025)
por: Mohammadi, Bahram, et al.
Publicado: (2025)
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
por: Cheng, Hongrong, et al.
Publicado: (2024)
por: Cheng, Hongrong, et al.
Publicado: (2024)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
por: Liu, Jiarui, et al.
Publicado: (2025)
por: Liu, Jiarui, et al.
Publicado: (2025)
Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics
por: Chrabąszcz, Maciej, et al.
Publicado: (2026)
por: Chrabąszcz, Maciej, et al.
Publicado: (2026)
What Makes a Representation Good for Single-Cell Perturbation Prediction?
por: Jiang, Wenkang, et al.
Publicado: (2026)
por: Jiang, Wenkang, et al.
Publicado: (2026)
ETAGE: Enhanced Test Time Adaptation with Integrated Entropy and Gradient Norms for Robust Model Performance
por: Shamsi, Afshar, et al.
Publicado: (2024)
por: Shamsi, Afshar, et al.
Publicado: (2024)
When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
por: Wang, Keyu, et al.
Publicado: (2025)
por: Wang, Keyu, et al.
Publicado: (2025)
Representational and Behavioral Stability of Truth in Large Language Models
por: Dies, Samantha, et al.
Publicado: (2025)
por: Dies, Samantha, et al.
Publicado: (2025)
CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation
por: Tong, Zhao, et al.
Publicado: (2026)
por: Tong, Zhao, et al.
Publicado: (2026)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
por: Li, Jiangnan, et al.
Publicado: (2025)
por: Li, Jiangnan, et al.
Publicado: (2025)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
por: Kirk, Hannah Rose, et al.
Publicado: (2024)
por: Kirk, Hannah Rose, et al.
Publicado: (2024)
Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone Ensembling
por: Rodriguez-Opazo, Cristian, et al.
Publicado: (2024)
por: Rodriguez-Opazo, Cristian, et al.
Publicado: (2024)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
por: Trinley, Katharina, et al.
Publicado: (2025)
por: Trinley, Katharina, et al.
Publicado: (2025)
PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Models
por: Zheng, Haoyu, et al.
Publicado: (2026)
por: Zheng, Haoyu, et al.
Publicado: (2026)
Do LLMs Know What Is Private Internally? Probing and Steering Contextual Privacy Norms in Large Language Model Representations
por: Wang, Haoran, et al.
Publicado: (2026)
por: Wang, Haoran, et al.
Publicado: (2026)
The Geometry of Tokens in Internal Representations of Large Language Models
por: Viswanathan, Karthik, et al.
Publicado: (2025)
por: Viswanathan, Karthik, et al.
Publicado: (2025)
Learning Latent Dynamical Causal Processes for Single-Cell Perturbation Prediction
por: Jiang, Wenkang, et al.
Publicado: (2026)
por: Jiang, Wenkang, et al.
Publicado: (2026)
Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoning
por: Liu, Shuo, et al.
Publicado: (2026)
por: Liu, Shuo, et al.
Publicado: (2026)
Probing Internal Representations of Multi-Word Verbs in Large Language Models
por: Kissane, Hassane, et al.
Publicado: (2025)
por: Kissane, Hassane, et al.
Publicado: (2025)
Precise Reasoning About Container-Internal Pointers with Logical Pinning
por: Guan, Yawen, et al.
Publicado: (2025)
por: Guan, Yawen, et al.
Publicado: (2025)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
por: Bhatia, Gagan, et al.
Publicado: (2026)
por: Bhatia, Gagan, et al.
Publicado: (2026)
Balancing Stylization and Truth via Disentangled Representation Steering
por: Shen, Chenglei, et al.
Publicado: (2025)
por: Shen, Chenglei, et al.
Publicado: (2025)
What Large Language Models Do Not Talk About: An Empirical Study of Moderation and Censorship Practices
por: Noels, Sander, et al.
Publicado: (2025)
por: Noels, Sander, et al.
Publicado: (2025)
When Names Disappear: Revealing What LLMs Actually Understand About Code
por: Le, Cuong Chi, et al.
Publicado: (2025)
por: Le, Cuong Chi, et al.
Publicado: (2025)
A Primer in Post-Training Reasoning Data: What We Know About How It Works
por: Li, Yaoming, et al.
Publicado: (2026)
por: Li, Yaoming, et al.
Publicado: (2026)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
por: Yamashita, Tomoya, et al.
Publicado: (2025)
por: Yamashita, Tomoya, et al.
Publicado: (2025)
InvariantStock: Learning Invariant Features for Mastering the Shifting Market
por: Cao, Haiyao, et al.
Publicado: (2024)
por: Cao, Haiyao, et al.
Publicado: (2024)
TInR: Exploring Tool-Internalized Reasoning in Large Language Models
por: Xu, Qiancheng, et al.
Publicado: (2026)
por: Xu, Qiancheng, et al.
Publicado: (2026)
Do LLMs Really Know What They Don't Know? Internal States Mainly Reflect Knowledge Recall Rather Than Truthfulness
por: Cheang, Chi Seng, et al.
Publicado: (2025)
por: Cheang, Chi Seng, et al.
Publicado: (2025)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
por: Ghosh, Shrestha, et al.
Publicado: (2025)
por: Ghosh, Shrestha, et al.
Publicado: (2025)
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
por: Mousavi, Seyed Mahed, et al.
Publicado: (2025)
por: Mousavi, Seyed Mahed, et al.
Publicado: (2025)
The Reasoning Error About Reasoning: Why Different Types of Reasoning Require Different Representational Structures
por: Wu, Yiling
Publicado: (2026)
por: Wu, Yiling
Publicado: (2026)
Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models
por: Song, Yusheng, et al.
Publicado: (2025)
por: Song, Yusheng, et al.
Publicado: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
por: Shamsi, Zafir, et al.
Publicado: (2026)
por: Shamsi, Zafir, et al.
Publicado: (2026)
Can Large Language Models Understand, Reason About, and Generate Code-Switched Text?
por: Winata, Genta Indra, et al.
Publicado: (2026)
por: Winata, Genta Indra, et al.
Publicado: (2026)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
por: Zhang, Shaolei, et al.
Publicado: (2024)
por: Zhang, Shaolei, et al.
Publicado: (2024)
Speed and Conversational Large Language Models: Not All Is About Tokens per Second
por: Conde, Javier, et al.
Publicado: (2025)
por: Conde, Javier, et al.
Publicado: (2025)
Ejemplares similares
-
Decomposing Task Vectors for Refined Model Editing
por: Damirchi, Hamed, et al.
Publicado: (2025) -
The Quest for Winning Tickets in Low-Rank Adapters
por: Damirchi, Hamed, et al.
Publicado: (2025) -
Semantic Role Labeling Guided Out-of-distribution Detection
por: Zou, Jinan, et al.
Publicado: (2023) -
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
por: Mohammadi, Bahram, et al.
Publicado: (2025) -
MINI-LLM: Memory-Efficient Structured Pruning for Large Language Models
por: Cheng, Hongrong, et al.
Publicado: (2024)