Mechanistic Indicators of Understanding in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Beckmann, Pierre, Queloz, Matthieu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
di: Queloz, Matthieu
Pubblicazione: (2025)
di: Queloz, Matthieu
Pubblicazione: (2025)
Probing Persona-Dependent Preferences in Language Models
di: Gilg, Oscar, et al.
Pubblicazione: (2026)
di: Gilg, Oscar, et al.
Pubblicazione: (2026)
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
di: Yu, Lei, et al.
Pubblicazione: (2024)
di: Yu, Lei, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Emotion Inference in Large Language Models
di: Tak, Ala N., et al.
Pubblicazione: (2025)
di: Tak, Ala N., et al.
Pubblicazione: (2025)
Mechanistic Decoding of Cognitive Constructs in Large Language Models
di: Shou, Yitong, et al.
Pubblicazione: (2026)
di: Shou, Yitong, et al.
Pubblicazione: (2026)
Where is the Mind? Persona Vectors and LLM Individuation
di: Beckmann, Pierre, et al.
Pubblicazione: (2026)
di: Beckmann, Pierre, et al.
Pubblicazione: (2026)
Detecting Linguistic Indicators for Stereotype Assessment with Large Language Models
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
di: Görge, Rebekka, et al.
Pubblicazione: (2025)
Mechanistic Decomposition of Sentence Representations
di: Tehenan, Matthieu, et al.
Pubblicazione: (2025)
di: Tehenan, Matthieu, et al.
Pubblicazione: (2025)
Entangled in Representations: Mechanistic Investigation of Cultural Biases in Large Language Models
di: Yu, Haeun, et al.
Pubblicazione: (2025)
di: Yu, Haeun, et al.
Pubblicazione: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
di: Wang, Shuxun, et al.
Pubblicazione: (2025)
Mechanistic Behavior Editing of Language Models
di: Singh, Joykirat, et al.
Pubblicazione: (2024)
di: Singh, Joykirat, et al.
Pubblicazione: (2024)
Potemkin Understanding in Large Language Models
di: Mancoridis, Marina, et al.
Pubblicazione: (2025)
di: Mancoridis, Marina, et al.
Pubblicazione: (2025)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
di: Chhabra, Vishnu Kabir, et al.
Pubblicazione: (2025)
Mechanistic Origin of Moral Indifference in Language Models
di: Li, Lingyu, et al.
Pubblicazione: (2026)
di: Li, Lingyu, et al.
Pubblicazione: (2026)
Emerging Opportunities of Using Large Language Models for Translation Between Drug Molecules and Indications
di: Oniani, David, et al.
Pubblicazione: (2024)
di: Oniani, David, et al.
Pubblicazione: (2024)
Understanding the Dilemma of Unlearning for Large Language Models
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
di: Zhang, Qingjie, et al.
Pubblicazione: (2025)
Evaluating Spatial Understanding of Large Language Models
di: Yamada, Yutaro, et al.
Pubblicazione: (2023)
di: Yamada, Yutaro, et al.
Pubblicazione: (2023)
Assessing and Understanding Creativity in Large Language Models
di: Zhao, Yunpu, et al.
Pubblicazione: (2024)
di: Zhao, Yunpu, et al.
Pubblicazione: (2024)
Toward Mechanistic Explanation of Deductive Reasoning in Language Models
di: Maltoni, Davide, et al.
Pubblicazione: (2025)
di: Maltoni, Davide, et al.
Pubblicazione: (2025)
Do Large Language Models Understand Word Senses?
di: Meconi, Domenico, et al.
Pubblicazione: (2025)
di: Meconi, Domenico, et al.
Pubblicazione: (2025)
Large Language Models Understanding: an Inherent Ambiguity Barrier
di: Nissani, Daniel N.
Pubblicazione: (2025)
di: Nissani, Daniel N.
Pubblicazione: (2025)
"Understanding AI": Semantic Grounding in Large Language Models
di: Lyre, Holger
Pubblicazione: (2024)
di: Lyre, Holger
Pubblicazione: (2024)
Metacognitive Prompting Improves Understanding in Large Language Models
di: Wang, Yuqing, et al.
Pubblicazione: (2023)
di: Wang, Yuqing, et al.
Pubblicazione: (2023)
Using Large Language Models to Understand Telecom Standards
di: Karapantelakis, Athanasios, et al.
Pubblicazione: (2024)
di: Karapantelakis, Athanasios, et al.
Pubblicazione: (2024)
Preservation of Language Understanding Capabilities in Speech-aware Large Language Models
di: Kubis, Marek, et al.
Pubblicazione: (2025)
di: Kubis, Marek, et al.
Pubblicazione: (2025)
Can Large Language Models Detect Verbal Indicators of Romantic Attraction?
di: Matz, Sandra C., et al.
Pubblicazione: (2024)
di: Matz, Sandra C., et al.
Pubblicazione: (2024)
Mechanistic Understanding of Language Models in Syntactic Code Completion
di: Miller, Samuel, et al.
Pubblicazione: (2025)
di: Miller, Samuel, et al.
Pubblicazione: (2025)
Modeling Understanding of Story-Based Analogies Using Large Language Models
di: Inani, Kalit, et al.
Pubblicazione: (2025)
di: Inani, Kalit, et al.
Pubblicazione: (2025)
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
di: Ball, Sarah, et al.
Pubblicazione: (2025)
di: Ball, Sarah, et al.
Pubblicazione: (2025)
Understanding Privacy Risks of Embeddings Induced by Large Language Models
di: Zhu, Zhihao, et al.
Pubblicazione: (2024)
di: Zhu, Zhihao, et al.
Pubblicazione: (2024)
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
di: Wu, Yulong, et al.
Pubblicazione: (2025)
di: Wu, Yulong, et al.
Pubblicazione: (2025)
Improving the Language Understanding Capabilities of Large Language Models Using Reinforcement Learning
di: Hu, Bokai, et al.
Pubblicazione: (2024)
di: Hu, Bokai, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of Socio-Political Frames in Language Models
di: Asghari, Hadi, et al.
Pubblicazione: (2025)
di: Asghari, Hadi, et al.
Pubblicazione: (2025)
Evaluating List Construction and Temporal Understanding capabilities of Large Language Models
di: Dumitru, Alexandru, et al.
Pubblicazione: (2025)
di: Dumitru, Alexandru, et al.
Pubblicazione: (2025)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
di: Kim, Jongho, et al.
Pubblicazione: (2025)
di: Kim, Jongho, et al.
Pubblicazione: (2025)
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
Understanding Data Temporality Impact on Large Language Models Pre-training
di: Pilchen, Hippolyte, et al.
Pubblicazione: (2026)
di: Pilchen, Hippolyte, et al.
Pubblicazione: (2026)
Can Large Language Models Understand Real-World Complex Instructions?
di: He, Qianyu, et al.
Pubblicazione: (2023)
di: He, Qianyu, et al.
Pubblicazione: (2023)
Enhancing Contextual Understanding in Large Language Models through Contrastive Decoding
di: Zhao, Zheng, et al.
Pubblicazione: (2024)
di: Zhao, Zheng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
di: Queloz, Matthieu
Pubblicazione: (2025) -
Probing Persona-Dependent Preferences in Language Models
di: Gilg, Oscar, et al.
Pubblicazione: (2026) -
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
di: Yu, Lei, et al.
Pubblicazione: (2024) -
Mechanistic Interpretability of Emotion Inference in Large Language Models
di: Tak, Ala N., et al.
Pubblicazione: (2025) -
Mechanistic Decoding of Cognitive Constructs in Large Language Models
di: Shou, Yitong, et al.
Pubblicazione: (2026)