Probing the Probes: Methods and Metrics for Concept Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Lysnæs-Larsen, Jacob, Eggen, Marte, Strümke, Inga |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Integrating attention into explanation frameworks for language and vision transformers
por: Eggen, Marte, et al.
Publicado: (2025)
por: Eggen, Marte, et al.
Publicado: (2025)
A transformer-based deep reinforcement learning approach to spatial navigation in a partially observable Morris Water Maze
por: Eggen, Marte, et al.
Publicado: (2024)
por: Eggen, Marte, et al.
Publicado: (2024)
A Framework for Causal Concept-based Model Explanations
por: Bjøru, Anna Rodum, et al.
Publicado: (2025)
por: Bjøru, Anna Rodum, et al.
Publicado: (2025)
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
por: Eggen, Marte, et al.
Publicado: (2026)
por: Eggen, Marte, et al.
Publicado: (2026)
Position: AI Security Policy Should Target Systems, Not Models
por: Riegler, Michael A., et al.
Publicado: (2026)
por: Riegler, Michael A., et al.
Publicado: (2026)
Evaluating Explainable AI Methods in Deep Learning Models for Early Detection of Cerebral Palsy
por: Pellano, Kimji N., et al.
Publicado: (2024)
por: Pellano, Kimji N., et al.
Publicado: (2024)
Choose Your Explanation: A Comparison of SHAP and GradCAM in Human Activity Recognition
por: Tempel, Felix, et al.
Publicado: (2024)
por: Tempel, Felix, et al.
Publicado: (2024)
Probing the Limits of Stylistic Alignment in Vision-Language Models
por: Farajidizaji, Asma, et al.
Publicado: (2025)
por: Farajidizaji, Asma, et al.
Publicado: (2025)
Polarity-Aware Probing for Quantifying Latent Alignment in Language Models
por: Sadiekh, Sabrina, et al.
Publicado: (2025)
por: Sadiekh, Sabrina, et al.
Publicado: (2025)
Towards Biomarker Discovery for Early Cerebral Palsy Detection: Evaluating Explanations Through Kinematic Perturbations
por: Pellano, Kimji N., et al.
Publicado: (2025)
por: Pellano, Kimji N., et al.
Publicado: (2025)
Probing the Emergence of Cross-lingual Alignment during LLM Training
por: Wang, Hetong, et al.
Publicado: (2024)
por: Wang, Hetong, et al.
Publicado: (2024)
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
por: Theodoridis, Nikos, et al.
Publicado: (2026)
por: Theodoridis, Nikos, et al.
Publicado: (2026)
Fidelity Probes for Specification--Code Alignment
por: Erata, Ferhat, et al.
Publicado: (2026)
por: Erata, Ferhat, et al.
Publicado: (2026)
Concept Probing: Where to Find Human-Defined Concepts (Extended Version)
por: Ribeiro, Manuel de Sousa, et al.
Publicado: (2025)
por: Ribeiro, Manuel de Sousa, et al.
Publicado: (2025)
On the Performance of Concept Probing: The Influence of the Data (Extended Version)
por: Ribeiro, Manuel de Sousa, et al.
Publicado: (2025)
por: Ribeiro, Manuel de Sousa, et al.
Publicado: (2025)
Does CLIP Bind Concepts? Probing Compositionality in Large Image Models
por: Lewis, Martha, et al.
Publicado: (2022)
por: Lewis, Martha, et al.
Publicado: (2022)
Probing for Consciousness in Machines
por: Immertreu, Mathis, et al.
Publicado: (2024)
por: Immertreu, Mathis, et al.
Publicado: (2024)
Latent Causal Probing: A Formal Perspective on Probing with Causal Models of Data
por: Jin, Charles, et al.
Publicado: (2024)
por: Jin, Charles, et al.
Publicado: (2024)
Concept Alignment
por: Rane, Sunayana, et al.
Publicado: (2024)
por: Rane, Sunayana, et al.
Publicado: (2024)
From Movements to Metrics: Evaluating Explainable AI Methods in Skeleton-Based Human Activity Recognition
por: Pellano, Kimji N., et al.
Publicado: (2024)
por: Pellano, Kimji N., et al.
Publicado: (2024)
Interplay between Federated Learning and Explainable Artificial Intelligence: a Scoping Review
por: Lopez-Ramos, Luis M., et al.
Publicado: (2024)
por: Lopez-Ramos, Luis M., et al.
Publicado: (2024)
The Geometry of Harmfulness in LLMs through Subconcept Probing
por: Shah, McNair, et al.
Publicado: (2025)
por: Shah, McNair, et al.
Publicado: (2025)
Linear Probe Penalties Reduce LLM Sycophancy
por: Papadatos, Henry, et al.
Publicado: (2024)
por: Papadatos, Henry, et al.
Publicado: (2024)
Deep Minds and Shallow Probes
por: Lee, Su Hyeong, et al.
Publicado: (2026)
por: Lee, Su Hyeong, et al.
Publicado: (2026)
Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design
por: Yang, Zaifei, et al.
Publicado: (2026)
por: Yang, Zaifei, et al.
Publicado: (2026)
Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities
por: Albright, Ryan, et al.
Publicado: (2026)
por: Albright, Ryan, et al.
Publicado: (2026)
Investigating Concept Alignment Using Implausible Category Members
por: Rane, Sunayana, et al.
Publicado: (2026)
por: Rane, Sunayana, et al.
Publicado: (2026)
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
por: Yang, Haoyun, et al.
Publicado: (2026)
por: Yang, Haoyun, et al.
Publicado: (2026)
Do Linear Probes Generalize Better in Persona Coordinates?
por: Mahadik, Prasad, et al.
Publicado: (2026)
por: Mahadik, Prasad, et al.
Publicado: (2026)
EL-VIT: Probing Vision Transformer with Interactive Visualization
por: Zhou, Hong, et al.
Publicado: (2024)
por: Zhou, Hong, et al.
Publicado: (2024)
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
por: Bai, Songlin, et al.
Publicado: (2026)
por: Bai, Songlin, et al.
Publicado: (2026)
Probe Pruning: Accelerating LLMs through Dynamic Pruning via Model-Probing
por: Le, Qi, et al.
Publicado: (2025)
por: Le, Qi, et al.
Publicado: (2025)
Probing Neural Combinatorial Optimization Models
por: Zhang, Zhiqin, et al.
Publicado: (2025)
por: Zhang, Zhiqin, et al.
Publicado: (2025)
Probing for Arithmetic Errors in Language Models
por: Sun, Yucheng, et al.
Publicado: (2025)
por: Sun, Yucheng, et al.
Publicado: (2025)
Probing Knowledge Holes in Unlearned LLMs
por: Ko, Myeongseob, et al.
Publicado: (2025)
por: Ko, Myeongseob, et al.
Publicado: (2025)
RAPTOR: Ridge-Adaptive Logistic Probes
por: Gao, Ziqi, et al.
Publicado: (2026)
por: Gao, Ziqi, et al.
Publicado: (2026)
Knowledge Probing for Graph Representation Learning
por: Zhao, Mingyu, et al.
Publicado: (2024)
por: Zhao, Mingyu, et al.
Publicado: (2024)
Probe-Geometry Alignment: Erasing the Cross-Sequence Memorization Signature Below Chance
por: Rupa, Anamika Paul, et al.
Publicado: (2026)
por: Rupa, Anamika Paul, et al.
Publicado: (2026)
Concept-Based Explainable Artificial Intelligence: Metrics and Benchmarks
por: Aysel, Halil Ibrahim, et al.
Publicado: (2025)
por: Aysel, Halil Ibrahim, et al.
Publicado: (2025)
GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMs
por: Eltahir, Mohamed, et al.
Publicado: (2026)
por: Eltahir, Mohamed, et al.
Publicado: (2026)
Ejemplares similares
-
Integrating attention into explanation frameworks for language and vision transformers
por: Eggen, Marte, et al.
Publicado: (2025) -
A transformer-based deep reinforcement learning approach to spatial navigation in a partially observable Morris Water Maze
por: Eggen, Marte, et al.
Publicado: (2024) -
A Framework for Causal Concept-based Model Explanations
por: Bjøru, Anna Rodum, et al.
Publicado: (2025) -
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
por: Eggen, Marte, et al.
Publicado: (2026) -
Position: AI Security Policy Should Target Systems, Not Models
por: Riegler, Michael A., et al.
Publicado: (2026)