Salvato in:
| Autori principali: | De Sabbata, Stef, Mizzaro, Stefano, Roitero, Kevin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.03368 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Geographic Relative Space in Large Language Models through Activation Patching
di: De Sabbata, Stef, et al.
Pubblicazione: (2026)
di: De Sabbata, Stef, et al.
Pubblicazione: (2026)
The Effect of Document Summarization on LLM-Based Relevance Judgments
di: Mohtadi, Samaneh, et al.
Pubblicazione: (2025)
di: Mohtadi, Samaneh, et al.
Pubblicazione: (2025)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
di: Lunardi, Riccardo, et al.
Pubblicazione: (2025)
di: Lunardi, Riccardo, et al.
Pubblicazione: (2025)
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
di: Roitero, Kevin, et al.
Pubblicazione: (2025)
di: Roitero, Kevin, et al.
Pubblicazione: (2025)
Rational Metareasoning for Large Language Models
di: De Sabbata, C. Nicolò, et al.
Pubblicazione: (2024)
di: De Sabbata, C. Nicolò, et al.
Pubblicazione: (2024)
GeoLLM: Extracting Geospatial Knowledge from Large Language Models
di: Manvi, Rohin, et al.
Pubblicazione: (2023)
di: Manvi, Rohin, et al.
Pubblicazione: (2023)
Binary Autoencoder for Mechanistic Interpretability of Large Language Models
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
di: Cho, Hakaze, et al.
Pubblicazione: (2025)
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
di: Winninger, Thomas, et al.
Pubblicazione: (2025)
di: Winninger, Thomas, et al.
Pubblicazione: (2025)
Mechanistic Interpretability with SAEs: Probing Religion, Violence, and Geography in Large Language Models
di: Simbeck, Katharina, et al.
Pubblicazione: (2025)
di: Simbeck, Katharina, et al.
Pubblicazione: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
Challenges in Mechanistically Interpreting Model Representations
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
di: Golechha, Satvik, et al.
Pubblicazione: (2024)
Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
di: García-Carrasco, Jorge, et al.
Pubblicazione: (2024)
Open Problems in Mechanistic Interpretability
di: Sharkey, Lee, et al.
Pubblicazione: (2025)
di: Sharkey, Lee, et al.
Pubblicazione: (2025)
Exemplar Partitioning for Mechanistic Interpretability
di: Rumbelow, Jessica
Pubblicazione: (2026)
di: Rumbelow, Jessica
Pubblicazione: (2026)
From Mechanistic to Compositional Interpretability
di: Gauderis, Ward, et al.
Pubblicazione: (2026)
di: Gauderis, Ward, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of RNNs emulating Hidden Markov Models
di: Torre, Elia, et al.
Pubblicazione: (2025)
di: Torre, Elia, et al.
Pubblicazione: (2025)
Reasoning Circuits in Language Models: A Mechanistic Interpretation of Syllogistic Inference
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
di: Kim, Geonhee, et al.
Pubblicazione: (2024)
Mechanistic Interpretability for Neural TSP Solvers
di: Narad, Reuben, et al.
Pubblicazione: (2025)
di: Narad, Reuben, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Reinforcement Learning Agents
di: Trim, Tristan, et al.
Pubblicazione: (2024)
di: Trim, Tristan, et al.
Pubblicazione: (2024)
Validating Mechanistic Interpretations: An Axiomatic Approach
di: Palumbo, Nils, et al.
Pubblicazione: (2024)
di: Palumbo, Nils, et al.
Pubblicazione: (2024)
Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
di: AlKhamissi, Badr, et al.
Pubblicazione: (2025)
di: AlKhamissi, Badr, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of LoRA-Adapted Language Models for Nuclear Reactor Safety Applications
di: Lee, Yoon Pyo
Pubblicazione: (2025)
di: Lee, Yoon Pyo
Pubblicazione: (2025)
Predicting missing values: A good idea?
di: van Buuren, Stef
Pubblicazione: (2026)
di: van Buuren, Stef
Pubblicazione: (2026)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
di: Wang, Xu, et al.
Pubblicazione: (2026)
di: Wang, Xu, et al.
Pubblicazione: (2026)
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
di: Baker, Mohammed Abu, et al.
Pubblicazione: (2025)
di: Baker, Mohammed Abu, et al.
Pubblicazione: (2025)
Using Degeneracy in the Loss Landscape for Mechanistic Interpretability
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
di: Bushnaq, Lucius, et al.
Pubblicazione: (2024)
Compact Proofs of Model Performance via Mechanistic Interpretability
di: Gross, Jason, et al.
Pubblicazione: (2024)
di: Gross, Jason, et al.
Pubblicazione: (2024)
Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
di: Wang, Shengyuan, et al.
Pubblicazione: (2025)
Mechanistic Interpretability Tool for AI Weather Models
di: Tempest, Kirsten I., et al.
Pubblicazione: (2026)
di: Tempest, Kirsten I., et al.
Pubblicazione: (2026)
Mechanistic Interpretability of Binary and Ternary Transformers
di: Li, Jason
Pubblicazione: (2024)
di: Li, Jason
Pubblicazione: (2024)
A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models
di: Lin, Zihao, et al.
Pubblicazione: (2025)
di: Lin, Zihao, et al.
Pubblicazione: (2025)
Mechanistic Interpretability of Brain-to-Speech Models Across Speech Modes
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
di: Maghsoudi, Maryam, et al.
Pubblicazione: (2026)
reward-lens: A Mechanistic Interpretability Library for Reward Models
di: Nadaf, Mohammed Suhail B
Pubblicazione: (2026)
di: Nadaf, Mohammed Suhail B
Pubblicazione: (2026)
TurnBack: A Geospatial Route Cognition Benchmark for Large Language Models through Reverse Route
di: Luo, Hongyi, et al.
Pubblicazione: (2025)
di: Luo, Hongyi, et al.
Pubblicazione: (2025)
Bridging Mechanistic Interpretability and Prompt Engineering with Gradient Ascent for Interpretable Persona Control
di: Saini, Harshvardhan, et al.
Pubblicazione: (2026)
di: Saini, Harshvardhan, et al.
Pubblicazione: (2026)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)
di: Mishra, Anurag
Pubblicazione: (2025)
OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting
di: Suri, Sanah, et al.
Pubblicazione: (2026)
di: Suri, Sanah, et al.
Pubblicazione: (2026)
LLMs as Implicit Imputers: Uncertainty Should Scale with Missing Information
di: van Buuren, Stef
Pubblicazione: (2026)
di: van Buuren, Stef
Pubblicazione: (2026)
Interpretable Deep Learning for Polar Mechanistic Reaction Prediction
di: Miller, Ryan J., et al.
Pubblicazione: (2025)
di: Miller, Ryan J., et al.
Pubblicazione: (2025)
Cluster-Based Random Forest Visualization and Interpretation
di: Sondag, Max, et al.
Pubblicazione: (2025)
di: Sondag, Max, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Exploring Geographic Relative Space in Large Language Models through Activation Patching
di: De Sabbata, Stef, et al.
Pubblicazione: (2026) -
The Effect of Document Summarization on LLM-Based Relevance Judgments
di: Mohtadi, Samaneh, et al.
Pubblicazione: (2025) -
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
di: Lunardi, Riccardo, et al.
Pubblicazione: (2025) -
Efficiency and Effectiveness of LLM-Based Summarization of Evidence in Crowdsourced Fact-Checking
di: Roitero, Kevin, et al.
Pubblicazione: (2025) -
Rational Metareasoning for Large Language Models
di: De Sabbata, C. Nicolò, et al.
Pubblicazione: (2024)