Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Michael, Subramani, Nishant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026)
by: Subramani, Nishant, et al.
Published: (2026)
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
by: Li, Michael, et al.
Published: (2026)
by: Li, Michael, et al.
Published: (2026)
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025)
by: Subramani, Nishant, et al.
Published: (2025)
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
by: Mutisya, Hillary, et al.
Published: (2026)
by: Mutisya, Hillary, et al.
Published: (2026)
Can Small Language Models Learn, Unlearn, and Retain Noise Patterns?
by: Scaria, Nicy, et al.
Published: (2024)
by: Scaria, Nicy, et al.
Published: (2024)
Multilingual Lexical Feature Analysis of Spoken Language for Predicting Major Depression Symptom Severity
by: Tokareva, Anastasiia, et al.
Published: (2025)
by: Tokareva, Anastasiia, et al.
Published: (2025)
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
by: Jacobi, Jonathan, et al.
Published: (2025)
by: Jacobi, Jonathan, et al.
Published: (2025)
The Geometry of Tokens in Internal Representations of Large Language Models
by: Viswanathan, Karthik, et al.
Published: (2025)
by: Viswanathan, Karthik, et al.
Published: (2025)
LLM Microscope: What Model Internals Reveal About Answer Correctness and Context Utilization
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
What Evidence Do Language Models Find Convincing?
by: Wan, Alexander, et al.
Published: (2024)
by: Wan, Alexander, et al.
Published: (2024)
Latent Feature Mining for Predictive Model Enhancement with Large Language Models
by: Li, Bingxuan, et al.
Published: (2024)
by: Li, Bingxuan, et al.
Published: (2024)
Can Perplexity Predict Fine-tuning Performance? An Investigation of Tokenization Effects on Sequential Language Models for Nepali
by: Luitel, Nishant, et al.
Published: (2024)
by: Luitel, Nishant, et al.
Published: (2024)
Sparse-Autoencoder-Guided Internal Representation Unlearning for Large Language Models
by: Yamashita, Tomoya, et al.
Published: (2025)
by: Yamashita, Tomoya, et al.
Published: (2025)
Attention Sinks as Internal Signals for Hallucination Detection in Large Language Models
by: Binkowski, Jakub, et al.
Published: (2026)
by: Binkowski, Jakub, et al.
Published: (2026)
On Lexical Invariance on Multisets and Graphs
by: Zhang, Muhan
Published: (2024)
by: Zhang, Muhan
Published: (2024)
TIDE: Textual Identity Detection for Evaluating and Augmenting Classification and Language Models
by: Klu, Emmanuel, et al.
Published: (2023)
by: Klu, Emmanuel, et al.
Published: (2023)
Transferring Linear Features Across Language Models With Model Stitching
by: Chen, Alan, et al.
Published: (2025)
by: Chen, Alan, et al.
Published: (2025)
Invariant Features in Language Models: Geometric Characterization and Model Attribution
by: Dasgupta, Agnibh, et al.
Published: (2026)
by: Dasgupta, Agnibh, et al.
Published: (2026)
Meta-Tuning LLMs to Leverage Lexical Knowledge for Generalizable Language Style Understanding
by: Guo, Ruohao, et al.
Published: (2023)
by: Guo, Ruohao, et al.
Published: (2023)
Mastering Board Games by External and Internal Planning with Language Models
by: Schultz, John, et al.
Published: (2024)
by: Schultz, John, et al.
Published: (2024)
Rotary Offset Features in Large Language Models
by: Jonasson, André
Published: (2025)
by: Jonasson, André
Published: (2025)
Patent Language Model Pretraining with ModernBERT
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
by: Yousefiramandi, Amirhossein, et al.
Published: (2025)
Transcoders Find Interpretable LLM Feature Circuits
by: Dunefsky, Jacob, et al.
Published: (2024)
by: Dunefsky, Jacob, et al.
Published: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
by: Wang, Xinyi, et al.
Published: (2023)
by: Wang, Xinyi, et al.
Published: (2023)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
by: Li, Zihao, et al.
Published: (2025)
by: Li, Zihao, et al.
Published: (2025)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Language Model Training Paradigms for Clinical Feature Embeddings
by: Hu, Yurong, et al.
Published: (2023)
by: Hu, Yurong, et al.
Published: (2023)
Semantic Structure of Feature Space in Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2026)
by: Kozlowski, Austin C., et al.
Published: (2026)
Truth as a Trajectory: What Internal Representations Reveal About Large Language Model Reasoning
by: Damirchi, Hamed, et al.
Published: (2026)
by: Damirchi, Hamed, et al.
Published: (2026)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
by: Juzek, Tom S., et al.
Published: (2024)
by: Juzek, Tom S., et al.
Published: (2024)
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs
by: Bethany, Mazal, et al.
Published: (2024)
by: Bethany, Mazal, et al.
Published: (2024)
Lexical Hints of Accuracy in LLM Reasoning Chains
by: Vanhoyweghen, Arne, et al.
Published: (2025)
by: Vanhoyweghen, Arne, et al.
Published: (2025)
Internal Planning in Language Models: Characterizing Horizon and Branch Awareness
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
by: Ustaomeroglu, Muhammed, et al.
Published: (2025)
KS-Lottery: Finding Certified Lottery Tickets for Multilingual Language Models
by: Yuan, Fei, et al.
Published: (2024)
by: Yuan, Fei, et al.
Published: (2024)
Automated Statistical Model Discovery with Language Models
by: Li, Michael Y., et al.
Published: (2024)
by: Li, Michael Y., et al.
Published: (2024)
Persistent Topological Features in Large Language Models
by: Gardinazzi, Yuri, et al.
Published: (2024)
by: Gardinazzi, Yuri, et al.
Published: (2024)
Finding Culture-Sensitive Neurons in Vision-Language Models
by: Zhao, Xiutian, et al.
Published: (2025)
by: Zhao, Xiutian, et al.
Published: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
by: Laptev, Daniil, et al.
Published: (2025)
by: Laptev, Daniil, et al.
Published: (2025)
Why is "Chicago" Predictive of Deceptive Reviews? Using LLMs to Discover Language Phenomena from Lexical Cues
by: Qu, Jiaming, et al.
Published: (2025)
by: Qu, Jiaming, et al.
Published: (2025)
Similar Items
-
Personal Information Parroting in Language Models
by: Subramani, Nishant, et al.
Published: (2026) -
MICE for CATs: Model-Internal Confidence Estimation for Calibrating Agents with Tools
by: Subramani, Nishant, et al.
Published: (2025) -
How Much Do Circuits Tell Us? Measuring the Consistency and Specificity of Language Model Circuits
by: Li, Michael, et al.
Published: (2026) -
SimBA: Simplifying Benchmark Analysis Using Performance Matrices Alone
by: Subramani, Nishant, et al.
Published: (2025) -
Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data
by: Mutisya, Hillary, et al.
Published: (2026)