Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Fuente:
arXiv
Saved in:
| Main Authors: | Kopf, Laura, Feldhus, Nils, Bykov, Kirill, Bommer, Philine Lou, Hedström, Anna, Höhne, Marina M. -C., Eberle, Oliver |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoSy: Evaluating Textual Explanations of Neurons
by: Kopf, Laura, et al.
Published: (2024)
by: Kopf, Laura, et al.
Published: (2024)
Interpreting Language Models Through Concept Descriptions: A Survey
by: Feldhus, Nils, et al.
Published: (2025)
by: Feldhus, Nils, et al.
Published: (2025)
Finding the right XAI method -- A Guide for the Evaluation and Ranking of Explainable AI Methods in Climate Science
by: Bommer, Philine, et al.
Published: (2023)
by: Bommer, Philine, et al.
Published: (2023)
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather
by: Bommer, Philine L., et al.
Published: (2025)
by: Bommer, Philine L., et al.
Published: (2025)
Labeling Neural Representations with Inverse Recognition
by: Bykov, Kirill, et al.
Published: (2023)
by: Bykov, Kirill, et al.
Published: (2023)
A Fresh Look at Sanity Checks for Saliency Maps
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
From Flexibility to Manipulation: The Slippery Slope of XAI Evaluation
by: Wickstrøm, Kristoffer, et al.
Published: (2024)
by: Wickstrøm, Kristoffer, et al.
Published: (2024)
Simplifying Outcomes of Language Model Component Analyses with ELIA
by: Eidt, Aaron Louis, et al.
Published: (2026)
by: Eidt, Aaron Louis, et al.
Published: (2026)
Evaluate with the Inverse: Efficient Approximation of Latent Explanation Quality Distribution
by: Eiras-Franco, Carlos, et al.
Published: (2025)
by: Eiras-Franco, Carlos, et al.
Published: (2025)
iFlip: Iterative Feedback-driven Counterfactual Example Refinement
by: Wang, Yilong, et al.
Published: (2026)
by: Wang, Yilong, et al.
Published: (2026)
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Free-text Rationale Generation under Readability Level Control
by: Hsu, Yi-Sheng, et al.
Published: (2024)
by: Hsu, Yi-Sheng, et al.
Published: (2024)
NeuronScope: A Multi-Agent Framework for Explaining Polysemantic Neurons in Language Models
by: Liu, Weiqi, et al.
Published: (2026)
by: Liu, Weiqi, et al.
Published: (2026)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
by: Hedström, Anna, et al.
Published: (2024)
by: Hedström, Anna, et al.
Published: (2024)
Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models
by: Möller, Sebastian, et al.
Published: (2025)
by: Möller, Sebastian, et al.
Published: (2025)
Explaining Text Similarity in Transformer Models
by: Vasileiou, Alexandros, et al.
Published: (2024)
by: Vasileiou, Alexandros, et al.
Published: (2024)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
by: Dhaini, Mahdi, et al.
Published: (2025)
by: Dhaini, Mahdi, et al.
Published: (2025)
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
by: Gupta, Ayush, et al.
Published: (2025)
by: Gupta, Ayush, et al.
Published: (2025)
Polysemanticity or Polysemy? Lexical Identity Confounds Superposition Metrics
by: Hou, Iyad Ait, et al.
Published: (2026)
by: Hou, Iyad Ait, et al.
Published: (2026)
Manipulating Feature Visualizations with Gradient Slingshots
by: Bareeva, Dilyara, et al.
Published: (2024)
by: Bareeva, Dilyara, et al.
Published: (2024)
PRISM: A Personality-Driven Multi-Agent Framework for Social Media Simulation
by: Lu, Zhixiang, et al.
Published: (2025)
by: Lu, Zhixiang, et al.
Published: (2025)
Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs
by: Ye, Charles, et al.
Published: (2026)
by: Ye, Charles, et al.
Published: (2026)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
by: Sun, Jingyi, et al.
Published: (2026)
by: Sun, Jingyi, et al.
Published: (2026)
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
by: Gong, Bofan, et al.
Published: (2025)
by: Gong, Bofan, et al.
Published: (2025)
Persona Prompting as a Lens on LLM Social Reasoning
by: Yang, Jing, et al.
Published: (2026)
by: Yang, Jing, et al.
Published: (2026)
Infherno: End-to-end Agent-based FHIR Resource Synthesis from Free-form Clinical Notes
by: Frei, Johann, et al.
Published: (2025)
by: Frei, Johann, et al.
Published: (2025)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments
by: Linhardt, Lorenz, et al.
Published: (2025)
by: Linhardt, Lorenz, et al.
Published: (2025)
PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
by: Loc, Ngoc Phan Phuoc, et al.
Published: (2026)
CoE: Chain-of-Explanation via Automatic Visual Concept Circuit Description and Polysemanticity Quantification
by: Yu, Wenlong, et al.
Published: (2025)
by: Yu, Wenlong, et al.
Published: (2025)
Multi-Intent Recognition in Dialogue Understanding: A Comparison Between Smaller Open-Source LLMs
by: Ahmad, Adnan, et al.
Published: (2025)
by: Ahmad, Adnan, et al.
Published: (2025)
PRISM: A Unified Framework for Post-Training LLMs Without Verifiable Rewards
by: Ghimire, Mukesh, et al.
Published: (2026)
by: Ghimire, Mukesh, et al.
Published: (2026)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
by: Wang, Qianli, et al.
Published: (2024)
by: Wang, Qianli, et al.
Published: (2024)
Trick or Neat: Adversarial Ambiguity and Language Model Evaluation
by: Karamolegkou, Antonia, et al.
Published: (2025)
by: Karamolegkou, Antonia, et al.
Published: (2025)
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations
by: Brandl, Stephanie, et al.
Published: (2024)
by: Brandl, Stephanie, et al.
Published: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering
by: Nahid, Md Mahadi Hasan, et al.
Published: (2025)
by: Nahid, Md Mahadi Hasan, et al.
Published: (2025)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
Explaining Bayesian Neural Networks
by: Bykov, Kirill, et al.
Published: (2021)
by: Bykov, Kirill, et al.
Published: (2021)
Similar Items
-
CoSy: Evaluating Textual Explanations of Neurons
by: Kopf, Laura, et al.
Published: (2024) -
Interpreting Language Models Through Concept Descriptions: A Survey
by: Feldhus, Nils, et al.
Published: (2025) -
Finding the right XAI method -- A Guide for the Evaluation and Ranking of Explainable AI Methods in Climate Science
by: Bommer, Philine, et al.
Published: (2023) -
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather
by: Bommer, Philine L., et al.
Published: (2025) -
Labeling Neural Representations with Inverse Recognition
by: Bykov, Kirill, et al.
Published: (2023)