Salvato in:
| Autori principali: | Lopardo, Gianluigi, Precioso, Frederic, Garreau, Damien |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.03485 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
Faithful and Robust Local Interpretability for Textual Predictions
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025)
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025)
Beyond Mixtures and Products for Ensemble Aggregation: A Likelihood Perspective on Generalized Means
di: Razafindralambo, Raphaël, et al.
Pubblicazione: (2026)
di: Razafindralambo, Raphaël, et al.
Pubblicazione: (2026)
Towards Understanding Steering Strength
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2026)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
When Are Two Scores Better Than One? Investigating Ensembles of Diffusion Models
di: Razafindralambo, Raphaël, et al.
Pubblicazione: (2026)
di: Razafindralambo, Raphaël, et al.
Pubblicazione: (2026)
Harnessing Large Language Models as Post-hoc Correctors
di: Zhong, Zhiqiang, et al.
Pubblicazione: (2024)
di: Zhong, Zhiqiang, et al.
Pubblicazione: (2024)
WolBanking77: Wolof Banking Speech Intent Classification Dataset
di: Kandji, Abdou Karim, et al.
Pubblicazione: (2025)
di: Kandji, Abdou Karim, et al.
Pubblicazione: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
di: Peng, Runyu, et al.
Pubblicazione: (2026)
di: Peng, Runyu, et al.
Pubblicazione: (2026)
A Symbolic Framework for Evaluating Mathematical Reasoning and Generalisation with Transformers
di: Meadows, Jordan, et al.
Pubblicazione: (2023)
di: Meadows, Jordan, et al.
Pubblicazione: (2023)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
di: Heyen, Henning, et al.
Pubblicazione: (2024)
di: Heyen, Henning, et al.
Pubblicazione: (2024)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
di: Dhaini, Mahdi, et al.
Pubblicazione: (2025)
Rethinking Expert Trajectory Utilization in LLM Post-training for Mathematical Reasoning
di: Ding, Bowen, et al.
Pubblicazione: (2025)
di: Ding, Bowen, et al.
Pubblicazione: (2025)
LLMs Explain't: A Post-Mortem on Semantic Interpretability in Transformer Models
di: Abdelhalim, Alhassan, et al.
Pubblicazione: (2026)
di: Abdelhalim, Alhassan, et al.
Pubblicazione: (2026)
CAM-Based Methods Can See through Walls
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2024)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2024)
The Risks of Recourse in Binary Classification
di: Fokkema, Hidde, et al.
Pubblicazione: (2023)
di: Fokkema, Hidde, et al.
Pubblicazione: (2023)
Feature Attribution from First Principles
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2025)
di: Taimeskhanov, Magamed, et al.
Pubblicazione: (2025)
On The Variability of Concept Activation Vectors
di: Wenkmann, Julia, et al.
Pubblicazione: (2025)
di: Wenkmann, Julia, et al.
Pubblicazione: (2025)
Attention Needs to Focus: A Unified Perspective on Attention Allocation
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
di: Fu, Zichuan, et al.
Pubblicazione: (2026)
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
di: Zhang, Zelin, et al.
Pubblicazione: (2026)
Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
di: Zhao, Dachuan, et al.
Pubblicazione: (2025)
SLM Meets LLM: Balancing Latency, Interpretability and Consistency in Hallucination Detection
di: Hu, Mengya, et al.
Pubblicazione: (2024)
di: Hu, Mengya, et al.
Pubblicazione: (2024)
Adaptive Two Sided Laplace Transforms: A Learnable, Interpretable, and Scalable Replacement for Self-Attention
di: Kiruluta, Andrew
Pubblicazione: (2025)
di: Kiruluta, Andrew
Pubblicazione: (2025)
Post-Training Sparse Attention with Double Sparsity
di: Yang, Shuo, et al.
Pubblicazione: (2024)
di: Yang, Shuo, et al.
Pubblicazione: (2024)
A Mathematical Philosophy of Explanations in Mechanistic Interpretability -- The Strange Science Part I.i
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
di: Ayonrinde, Kola, et al.
Pubblicazione: (2025)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
di: Xian, Ruicheng, et al.
Pubblicazione: (2025)
Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language Models
di: kurra, Sailesh kiran, et al.
Pubblicazione: (2026)
di: kurra, Sailesh kiran, et al.
Pubblicazione: (2026)
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities
di: Sun, Hao, et al.
Pubblicazione: (2025)
di: Sun, Hao, et al.
Pubblicazione: (2025)
Improving Prediction Performance and Model Interpretability through Attention Mechanisms from Basic and Applied Research Perspectives
di: Kitada, Shunsuke
Pubblicazione: (2023)
di: Kitada, Shunsuke
Pubblicazione: (2023)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
di: Neo, Clement, et al.
Pubblicazione: (2024)
di: Neo, Clement, et al.
Pubblicazione: (2024)
Mind the map! Accounting for existing map information when estimating online HDMaps from sensor
di: Sun, Rémy, et al.
Pubblicazione: (2023)
di: Sun, Rémy, et al.
Pubblicazione: (2023)
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective
di: Bai, Xueying, et al.
Pubblicazione: (2024)
di: Bai, Xueying, et al.
Pubblicazione: (2024)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
di: He, Shenghua, et al.
Pubblicazione: (2025)
di: He, Shenghua, et al.
Pubblicazione: (2025)
How do Large Language Models Understand Relevance? A Mechanistic Interpretability Perspective
di: Liu, Qi, et al.
Pubblicazione: (2025)
di: Liu, Qi, et al.
Pubblicazione: (2025)
A High-Resolution Landscape Dataset for Concept-Based XAI With Application to Species Distribution Models
di: de la Brosse, Augustin, et al.
Pubblicazione: (2026)
di: de la Brosse, Augustin, et al.
Pubblicazione: (2026)
Detecting Suicidal Ideation in Text with Interpretable Deep Learning: A CNN-BiGRU with Attention Mechanism
di: Bhuiyan, Mohaiminul Islam, et al.
Pubblicazione: (2025)
di: Bhuiyan, Mohaiminul Islam, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Understanding Post-hoc Explainers: The Case of Anchors
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023) -
Faithful and Robust Local Interpretability for Textual Predictions
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2023) -
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022) -
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022) -
MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations
di: Mitsuzawa, Kensuke, et al.
Pubblicazione: (2025)