Mechanistic Indicators of Understanding in Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Beckmann, Pierre, Queloz, Matthieu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
par: Queloz, Matthieu
Publié: (2025)
par: Queloz, Matthieu
Publié: (2025)
Probing Persona-Dependent Preferences in Language Models
par: Gilg, Oscar, et autres
Publié: (2026)
par: Gilg, Oscar, et autres
Publié: (2026)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
par: Yu, Lei, et autres
Publié: (2024)
par: Yu, Lei, et autres
Publié: (2024)
Mechanistic Interpretability of Emotion Inference in Large Language Models
par: Tak, Ala N., et autres
Publié: (2025)
par: Tak, Ala N., et autres
Publié: (2025)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
par: Shou, Yitong, et autres
Publié: (2026)
par: Shou, Yitong, et autres
Publié: (2026)
Where is the Mind? Persona Vectors and LLM Individuation
par: Beckmann, Pierre, et autres
Publié: (2026)
par: Beckmann, Pierre, et autres
Publié: (2026)
Detecting Linguistic Indicators for Stereotype Assessment with Large Language Models
par: Görge, Rebekka, et autres
Publié: (2025)
par: Görge, Rebekka, et autres
Publié: (2025)
Mechanistic Decomposition of Sentence Representations
par: Tehenan, Matthieu, et autres
Publié: (2025)
par: Tehenan, Matthieu, et autres
Publié: (2025)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
par: Yu, Haeun, et autres
Publié: (2025)
par: Yu, Haeun, et autres
Publié: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
par: Wang, Shuxun, et autres
Publié: (2025)
par: Wang, Shuxun, et autres
Publié: (2025)
Mechanistic Behavior Editing of Language Models
par: Singh, Joykirat, et autres
Publié: (2024)
par: Singh, Joykirat, et autres
Publié: (2024)
Potemkin Understanding in Large Language Models
par: Mancoridis, Marina, et autres
Publié: (2025)
par: Mancoridis, Marina, et autres
Publié: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
par: Chhabra, Vishnu Kabir, et autres
Publié: (2025)
par: Chhabra, Vishnu Kabir, et autres
Publié: (2025)
Mechanistic Origin of Moral Indifference in Language Models
par: Li, Lingyu, et autres
Publié: (2026)
par: Li, Lingyu, et autres
Publié: (2026)
Emerging Opportunities of Using Large Language Models for Translation Between Drug Molecules and Indications
par: Oniani, David, et autres
Publié: (2024)
par: Oniani, David, et autres
Publié: (2024)
Understanding the Dilemma of Unlearning for Large Language Models
par: Zhang, Qingjie, et autres
Publié: (2025)
par: Zhang, Qingjie, et autres
Publié: (2025)
Evaluating Spatial Understanding of Large Language Models
par: Yamada, Yutaro, et autres
Publié: (2023)
par: Yamada, Yutaro, et autres
Publié: (2023)
Assessing and Understanding Creativity in Large Language Models
par: Zhao, Yunpu, et autres
Publié: (2024)
par: Zhao, Yunpu, et autres
Publié: (2024)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
par: Maltoni, Davide, et autres
Publié: (2025)
par: Maltoni, Davide, et autres
Publié: (2025)
Do Large Language Models Understand Word Senses?
par: Meconi, Domenico, et autres
Publié: (2025)
par: Meconi, Domenico, et autres
Publié: (2025)
Large Language Models Understanding: an Inherent Ambiguity Barrier
par: Nissani, Daniel N.
Publié: (2025)
par: Nissani, Daniel N.
Publié: (2025)
"Understanding AI": Semantic Grounding in Large Language Models
par: Lyre, Holger
Publié: (2024)
par: Lyre, Holger
Publié: (2024)
Metacognitive Prompting Improves Understanding in Large Language Models
par: Wang, Yuqing, et autres
Publié: (2023)
par: Wang, Yuqing, et autres
Publié: (2023)
Using Large Language Models to Understand Telecom Standards
par: Karapantelakis, Athanasios, et autres
Publié: (2024)
par: Karapantelakis, Athanasios, et autres
Publié: (2024)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
par: Kubis, Marek, et autres
Publié: (2025)
par: Kubis, Marek, et autres
Publié: (2025)
Can Large Language Models Detect Verbal Indicators of Romantic Attraction?
par: Matz, Sandra C., et autres
Publié: (2024)
par: Matz, Sandra C., et autres
Publié: (2024)
Mechanistic Understanding of Language Models in Syntactic Code Completion
par: Miller, Samuel, et autres
Publié: (2025)
par: Miller, Samuel, et autres
Publié: (2025)
Modeling Understanding of Story-Based Analogies Using Large Language Models
par: Inani, Kalit, et autres
Publié: (2025)
par: Inani, Kalit, et autres
Publié: (2025)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
par: Ball, Sarah, et autres
Publié: (2025)
par: Ball, Sarah, et autres
Publié: (2025)
Understanding Privacy Risks of Embeddings Induced by Large Language Models
par: Zhu, Zhihao, et autres
Publié: (2024)
par: Zhu, Zhihao, et autres
Publié: (2024)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
par: Wu, Yulong, et autres
Publié: (2025)
par: Wu, Yulong, et autres
Publié: (2025)
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning
par: Hu, Bokai, et autres
Publié: (2024)
par: Hu, Bokai, et autres
Publié: (2024)
Mechanistic Interpretability of Socio-Political Frames in Language Models
par: Asghari, Hadi, et autres
Publié: (2025)
par: Asghari, Hadi, et autres
Publié: (2025)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
par: Dumitru, Alexandru, et autres
Publié: (2025)
par: Dumitru, Alexandru, et autres
Publié: (2025)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
par: Kim, Jongho, et autres
Publié: (2025)
par: Kim, Jongho, et autres
Publié: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
par: Zhang, Taolin, et autres
Publié: (2025)
par: Zhang, Taolin, et autres
Publié: (2025)
Understanding Data Temporality Impact on Large Language Models Pre-training
par: Pilchen, Hippolyte, et autres
Publié: (2026)
par: Pilchen, Hippolyte, et autres
Publié: (2026)
Can Large Language Models Understand Real-World Complex Instructions?
par: He, Qianyu, et autres
Publié: (2023)
par: He, Qianyu, et autres
Publié: (2023)
Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
par: Zhao, Zheng, et autres
Publié: (2024)
par: Zhao, Zheng, et autres
Publié: (2024)
Documents similaires
-
Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
par: Queloz, Matthieu
Publié: (2025) -
Probing Persona-Dependent Preferences in Language Models
par: Gilg, Oscar, et autres
Publié: (2026) -
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
par: Yu, Lei, et autres
Publié: (2024) -
Mechanistic Interpretability of Emotion Inference in Large Language Models
par: Tak, Ala N., et autres
Publié: (2025) -
Mechanistic Decoding of Cognitive Constructs in Large Language Models
par: Shou, Yitong, et autres
Publié: (2026)