Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yukun, Dong, Qi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917008366370816
author Zhang, Yukun
Dong, Qi
author_facet Zhang, Yukun
Dong, Qi
contents We present Multi-Scale Manifold Alignment(MSMA), an information-geometric framework that decomposes LLM representations into local, intermediate, and global manifolds and learns cross-scale mappings that preserve geometry and information. Across GPT-2, BERT, RoBERTa, and T5, we observe consistent hierarchical patterns and find that MSMA improves alignment metrics under multiple estimators (e.g., relative KL reduction and MI gains with statistical significance across seeds). Controlled interventions at different scales yield distinct and architecture-dependent effects on lexical diversity, sentence structure, and discourse coherence. While our theoretical analysis relies on idealized assumptions, the empirical results suggest that multi-objective alignment offers a practical lens for analyzing cross-scale information flow and guiding representation-level control.
format Preprint
id arxiv_https___arxiv_org_abs_2505_20333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
Zhang, Yukun
Dong, Qi
Computation and Language
Artificial Intelligence
We present Multi-Scale Manifold Alignment(MSMA), an information-geometric framework that decomposes LLM representations into local, intermediate, and global manifolds and learns cross-scale mappings that preserve geometry and information. Across GPT-2, BERT, RoBERTa, and T5, we observe consistent hierarchical patterns and find that MSMA improves alignment metrics under multiple estimators (e.g., relative KL reduction and MI gains with statistical significance across seeds). Controlled interventions at different scales yield distinct and architecture-dependent effects on lexical diversity, sentence structure, and discourse coherence. While our theoretical analysis relies on idealized assumptions, the empirical results suggest that multi-objective alignment offers a practical lens for analyzing cross-scale information flow and guiding representation-level control.
title Multi-Scale Manifold Alignment for Interpreting Large Language Models: A Unified Information-Geometric Framework
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.20333