Geospatial Mechanistic Interpretability of Large Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | De Sabbata, Stef, Mizzaro, Stefano, Roitero, Kevin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Exploring Geographic Relative Space in Large Language Models through Activation Patching
par: De Sabbata, Stef, et autres
Publié: (2026)
par: De Sabbata, Stef, et autres
Publié: (2026)
The Effect of Document Summarization on LLM-Based Relevance Judgments
par: Mohtadi, Samaneh, et autres
Publié: (2025)
par: Mohtadi, Samaneh, et autres
Publié: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
par: Lunardi, Riccardo, et autres
Publié: (2025)
par: Lunardi, Riccardo, et autres
Publié: (2025)
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
par: Roitero, Kevin, et autres
Publié: (2025)
par: Roitero, Kevin, et autres
Publié: (2025)
Rational Metareasoning for Large Language Models
par: De Sabbata, C. Nicolò, et autres
Publié: (2024)
par: De Sabbata, C. Nicolò, et autres
Publié: (2024)
GeoLLM: Extracting Geospatial Knowledge from Large Language Models
par: Manvi, Rohin, et autres
Publié: (2023)
par: Manvi, Rohin, et autres
Publié: (2023)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
par: Cho, Hakaze, et autres
Publié: (2025)
par: Cho, Hakaze, et autres
Publié: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
par: Winninger, Thomas, et autres
Publié: (2025)
par: Winninger, Thomas, et autres
Publié: (2025)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
par: Simbeck, Katharina, et autres
Publié: (2025)
par: Simbeck, Katharina, et autres
Publié: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
par: Liu, Qi, et autres
Publié: (2025)
par: Liu, Qi, et autres
Publié: (2025)
Challenges in Mechanistically Interpreting Model Representations
par: Golechha, Satvik, et autres
Publié: (2024)
par: Golechha, Satvik, et autres
Publié: (2024)
Open Problems in Mechanistic Interpretability
par: Sharkey, Lee, et autres
Publié: (2025)
par: Sharkey, Lee, et autres
Publié: (2025)
Exemplar Partitioning for Mechanistic Interpretability
par: Rumbelow, Jessica
Publié: (2026)
par: Rumbelow, Jessica
Publié: (2026)
From Mechanistic to Compositional Interpretability
par: Gauderis, Ward, et autres
Publié: (2026)
par: Gauderis, Ward, et autres
Publié: (2026)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
par: Torre, Elia, et autres
Publié: (2025)
par: Torre, Elia, et autres
Publié: (2025)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
par: García-Carrasco, Jorge, et autres
Publié: (2024)
par: García-Carrasco, Jorge, et autres
Publié: (2024)
Mechanistic Interpretability for Neural TSP Solvers
par: Narad, Reuben, et autres
Publié: (2025)
par: Narad, Reuben, et autres
Publié: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
par: Trim, Tristan, et autres
Publié: (2024)
par: Trim, Tristan, et autres
Publié: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
par: Palumbo, Nils, et autres
Publié: (2024)
par: Palumbo, Nils, et autres
Publié: (2024)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
par: Kim, Geonhee, et autres
Publié: (2024)
par: Kim, Geonhee, et autres
Publié: (2024)
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
par: Lee, Yoon Pyo
Publié: (2025)
par: Lee, Yoon Pyo
Publié: (2025)
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
par: Baker, Mohammed Abu, et autres
Publié: (2025)
par: Baker, Mohammed Abu, et autres
Publié: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
par: Bushnaq, Lucius, et autres
Publié: (2024)
par: Bushnaq, Lucius, et autres
Publié: (2024)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
par: Wang, Xu, et autres
Publié: (2026)
par: Wang, Xu, et autres
Publié: (2026)
Predicting missing values: A good idea?
par: van Buuren, Stef
Publié: (2026)
par: van Buuren, Stef
Publié: (2026)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
par: Saini, Harshvardhan, et autres
Publié: (2026)
par: Saini, Harshvardhan, et autres
Publié: (2026)
OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting
par: Suri, Sanah, et autres
Publié: (2026)
par: Suri, Sanah, et autres
Publié: (2026)
Compact Proofs of Model Performance via Mechanistic Interpretability
par: Gross, Jason, et autres
Publié: (2024)
par: Gross, Jason, et autres
Publié: (2024)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
par: Wang, Shengyuan, et autres
Publié: (2025)
par: Wang, Shengyuan, et autres
Publié: (2025)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
par: Lin, Zihao, et autres
Publié: (2025)
par: Lin, Zihao, et autres
Publié: (2025)
TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
par: Luo, Hongyi, et autres
Publié: (2025)
par: Luo, Hongyi, et autres
Publié: (2025)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
par: Maghsoudi, Maryam, et autres
Publié: (2026)
par: Maghsoudi, Maryam, et autres
Publié: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
par: Nadaf, Mohammed Suhail B
Publié: (2026)
par: Nadaf, Mohammed Suhail B
Publié: (2026)
Route Sparse Autoencoder to Interpret Large Language Models
par: Shi, Wei, et autres
Publié: (2025)
par: Shi, Wei, et autres
Publié: (2025)
Interpreting Learned Feedback Patterns in Large Language Models
par: Marks, Luke, et autres
Publié: (2023)
par: Marks, Luke, et autres
Publié: (2023)
Mechanistic Interpretability of Binary and Ternary Transformers
par: Li, Jason
Publié: (2024)
par: Li, Jason
Publié: (2024)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
par: Miller, Ryan J., et autres
Publié: (2025)
par: Miller, Ryan J., et autres
Publié: (2025)
Mechanistic Interpretability Tool for AI Weather Models
par: Tempest, Kirsten I., et autres
Publié: (2026)
par: Tempest, Kirsten I., et autres
Publié: (2026)
Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
par: AlKhamissi, Badr, et autres
Publié: (2025)
par: AlKhamissi, Badr, et autres
Publié: (2025)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
par: Mishra, Anurag
Publié: (2025)
par: Mishra, Anurag
Publié: (2025)
Documents similaires
-
Exploring Geographic Relative Space in Large Language Models through Activation Patching
par: De Sabbata, Stef, et autres
Publié: (2026) -
The Effect of Document Summarization on LLM-Based Relevance Judgments
par: Mohtadi, Samaneh, et autres
Publié: (2025) -
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
par: Lunardi, Riccardo, et autres
Publié: (2025) -
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
par: Roitero, Kevin, et autres
Publié: (2025) -
Rational Metareasoning for Large Language Models
par: De Sabbata, C. Nicolò, et autres
Publié: (2024)