LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Jiarui, Jain, Jivitesh, Diab, Mona, Subramani, Nishant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
by: Li, Michael, et al.
Published: (2025)
by: Li, Michael, et al.
Published: (2025)
Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
by: Liu, Jiarui, et al.
Published: (2024)
by: Liu, Jiarui, et al.
Published: (2024)
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design
by: Xiao, Yunze, et al.
Published: (2025)
by: Xiao, Yunze, et al.
Published: (2025)
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
by: Li, Michael, et al.
Published: (2026)
by: Li, Michael, et al.
Published: (2026)
BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data
by: Li, Wenkai, et al.
Published: (2024)
by: Li, Wenkai, et al.
Published: (2024)
Taming Object Hallucinations with Verified Atomic Confidence Estimation
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Generative Value Conflicts Reveal LLM Priorities
by: Liu, Andy, et al.
Published: (2025)
by: Liu, Andy, et al.
Published: (2025)
Evaluating Large Language Model Biases in Persona-Steered Generation
by: Liu, Andy, et al.
Published: (2024)
by: Liu, Andy, et al.
Published: (2024)
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
by: Damirchi, Hamed, et al.
Published: (2026)
by: Damirchi, Hamed, et al.
Published: (2026)
A Note on Bias to Complete
by: Xu, Jia, et al.
Published: (2024)
by: Xu, Jia, et al.
Published: (2024)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
NLI under the Microscope: What Atomic Hypothesis Decomposition Reveals
by: Srikanth, Neha, et al.
Published: (2025)
by: Srikanth, Neha, et al.
Published: (2025)
Combining Discrete Wavelet and Cosine Transforms for Efficient Sentence Embedding
by: Salama, Rana, et al.
Published: (2025)
by: Salama, Rana, et al.
Published: (2025)
Mining the Mind: What 100M Beliefs Reveal About Frontier LLM Knowledge
by: Ghosh, Shrestha, et al.
Published: (2025)
by: Ghosh, Shrestha, et al.
Published: (2025)
Defragmenting Language Models: An Interpretability-based Approach for Vocabulary Expansion
by: Mehta, Maitrey, et al.
Published: (2026)
by: Mehta, Maitrey, et al.
Published: (2026)
EVALUESTEER: Measuring Reward Model Steerability Towards Values and Preferences
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes
by: Ghate, Kshitish, et al.
Published: (2025)
by: Ghate, Kshitish, et al.
Published: (2025)
Decoding Dark Matter: Specialized Sparse Autoencoders for Interpreting Rare Concepts in Foundation Models
by: Muhamed, Aashiq, et al.
Published: (2024)
by: Muhamed, Aashiq, et al.
Published: (2024)
Semantic Compression for Word and Sentence Embeddings using Discrete Wavelet Transform
by: Salama, Rana Aref, et al.
Published: (2025)
by: Salama, Rana Aref, et al.
Published: (2025)
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
by: Liu, Jiarui, et al.
Published: (2026)
by: Liu, Jiarui, et al.
Published: (2026)
Towards Global AI Inclusivity: A Large-Scale Multilingual Terminology Dataset (GIST)
by: Liu, Jiarui, et al.
Published: (2024)
by: Liu, Jiarui, et al.
Published: (2024)
StressRoBERTa: Cross-Condition Transfer Learning from Depression, Anxiety, and PTSD to Stress Detection
by: Alqahtani, Amal, et al.
Published: (2025)
by: Alqahtani, Amal, et al.
Published: (2025)
DWTSumm: Discrete Wavelet Transform for Document Summarization
by: Salama, Rana, et al.
Published: (2026)
by: Salama, Rana, et al.
Published: (2026)
Evaluating and Calibrating LLM Confidence on Questions with Multiple Correct Answers
by: Wang, Yuhan, et al.
Published: (2026)
by: Wang, Yuhan, et al.
Published: (2026)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can't Answer?
by: Balepur, Nishant, et al.
Published: (2024)
by: Balepur, Nishant, et al.
Published: (2024)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
Investigating Cultural Alignment of Large Language Models
by: AlKhamissi, Badr, et al.
Published: (2024)
by: AlKhamissi, Badr, et al.
Published: (2024)
LLM Internal States Reveal Hallucination Risk Faced With a Query
by: Ji, Ziwei, et al.
Published: (2024)
by: Ji, Ziwei, et al.
Published: (2024)
See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses
by: Chen, Yulong, et al.
Published: (2024)
by: Chen, Yulong, et al.
Published: (2024)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
by: Sauter, Adrian, et al.
Published: (2026)
by: Sauter, Adrian, et al.
Published: (2026)
When Names Disappear: Revealing What LLMs Actually Understand About Code
by: Le, Cuong Chi, et al.
Published: (2025)
by: Le, Cuong Chi, et al.
Published: (2025)
Make Your LLM Fully Utilize the Context
by: An, Shengnan, et al.
Published: (2024)
by: An, Shengnan, et al.
Published: (2024)
Depth-Wise Attention (DWAtt): A Layer Fusion Method for Data-Efficient Classification
by: ElNokrashy, Muhammad, et al.
Published: (2022)
by: ElNokrashy, Muhammad, et al.
Published: (2022)
LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers
by: Razzhigaev, Anton, et al.
Published: (2025)
by: Razzhigaev, Anton, et al.
Published: (2025)
Similar Items
-
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026) -
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025) -
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
by: Li, Michael, et al.
Published: (2025) -
Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
by: Liu, Jiarui, et al.
Published: (2024) -
Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of Design
by: Xiao, Yunze, et al.
Published: (2025)