Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dorszewski, Teresa, Tětková, Lenka, Linhardt, Lorenz, Hansen, Lars Kai
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908047827271680
author Dorszewski, Teresa
Tětková, Lenka
Linhardt, Lorenz
Hansen, Lars Kai
author_facet Dorszewski, Teresa
Tětková, Lenka
Linhardt, Lorenz
Hansen, Lars Kai
contents Understanding how neural networks align with human cognitive processes is a crucial step toward developing more interpretable and reliable AI systems. Motivated by theories of human cognition, this study examines the relationship between \emph{convexity} in neural network representations and \emph{human-machine alignment} based on behavioral data. We identify a correlation between these two dimensions in pretrained and fine-tuned vision transformer models. Our findings suggest that the convex regions formed in latent spaces of neural networks to some extent align with human-defined categories and reflect the similarity relations humans use in cognitive tasks. While optimizing for alignment generally enhances convexity, increasing convexity through fine-tuning yields inconsistent effects on alignment, which suggests a complex relationship between the two. This study presents a first step toward understanding the relationship between the convexity of latent representations and human-machine alignment.
format Preprint
id arxiv_https___arxiv_org_abs_2409_06362
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks
Dorszewski, Teresa
Tětková, Lenka
Linhardt, Lorenz
Hansen, Lars Kai
Machine Learning
Artificial Intelligence
Understanding how neural networks align with human cognitive processes is a crucial step toward developing more interpretable and reliable AI systems. Motivated by theories of human cognition, this study examines the relationship between \emph{convexity} in neural network representations and \emph{human-machine alignment} based on behavioral data. We identify a correlation between these two dimensions in pretrained and fine-tuned vision transformer models. Our findings suggest that the convex regions formed in latent spaces of neural networks to some extent align with human-defined categories and reflect the similarity relations humans use in cognitive tasks. While optimizing for alignment generally enhances convexity, increasing convexity through fine-tuning yields inconsistent effects on alignment, which suggests a complex relationship between the two. This study presents a first step toward understanding the relationship between the convexity of latent representations and human-machine alignment.
title Connecting Concept Convexity and Human-Machine Alignment in Deep Neural Networks
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.06362