GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Pan, Tianyu Bell, Woodard, Damon L.
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908901907103744
author Pan, Tianyu Bell
Woodard, Damon L.
author_facet Pan, Tianyu Bell
Woodard, Damon L.
contents Large language models (LLMs) demonstrate strong performance, but they often lack transparency. We introduce GeoLAN, a training framework that treats token representations as geometric trajectories and applies stickiness conditions inspired by recent developments related to the Kakeya Conjecture. We have developed two differentiable regularizers, Katz-Tao Convex Wolff (KT-CW) and Katz-Tao Attention (KT-Attn), that promote isotropy and encourage diverse attention. Our experiments with Gemma-3 (1B, 4B, 12B) and Llama-3-8B show that GeoLAN frequently maintains task accuracy while improving geometric metrics and reducing certain fairness biases. These benefits are most significant in mid-sized models. Our findings reveal scale-dependent trade-offs between geometric precision and performance, suggesting that geometry-aware training is a promising approach to enhance mechanistic interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19460
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models
Pan, Tianyu Bell
Woodard, Damon L.
Machine Learning
Computational Geometry
Large language models (LLMs) demonstrate strong performance, but they often lack transparency. We introduce GeoLAN, a training framework that treats token representations as geometric trajectories and applies stickiness conditions inspired by recent developments related to the Kakeya Conjecture. We have developed two differentiable regularizers, Katz-Tao Convex Wolff (KT-CW) and Katz-Tao Attention (KT-Attn), that promote isotropy and encourage diverse attention. Our experiments with Gemma-3 (1B, 4B, 12B) and Llama-3-8B show that GeoLAN frequently maintains task accuracy while improving geometric metrics and reducing certain fairness biases. These benefits are most significant in mid-sized models. Our findings reveal scale-dependent trade-offs between geometric precision and performance, suggesting that geometry-aware training is a promising approach to enhance mechanistic interpretability.
title GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models
topic Machine Learning
Computational Geometry
url https://arxiv.org/abs/2603.19460