CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Zihao, Liao, Ming, Russell, Chris, Cai, Zhenguang G. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
by: Fu, Zihao, et al.
Published: (2026)
by: Fu, Zihao, et al.
Published: (2026)
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Understanding Transformers via N-gram Statistics
by: Nguyen, Timothy
Published: (2024)
by: Nguyen, Timothy
Published: (2024)
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025)
by: Yao, Xinhao, et al.
Published: (2025)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Calibration Across Layers: Understanding Calibration Evolution in LLMs
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
by: Lugoloobi, William, et al.
Published: (2026)
by: Lugoloobi, William, et al.
Published: (2026)
Understanding Subword Compositionality of Large Language Models
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
by: Kim, Jeonghoon, et al.
Published: (2025)
by: Kim, Jeonghoon, et al.
Published: (2025)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Understanding Dynamic Compute Allocation in Recurrent Transformers
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
by: Moosa, Ibraheem Muhammad, et al.
Published: (2026)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
by: Yehudai, Gilad, et al.
Published: (2025)
by: Yehudai, Gilad, et al.
Published: (2025)
Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
by: Tang, Xuemei, et al.
Published: (2024)
by: Tang, Xuemei, et al.
Published: (2024)
Diversity of Transformer Layers: One Aspect of Parameter Scaling Laws
by: Kamigaito, Hidetaka, et al.
Published: (2025)
by: Kamigaito, Hidetaka, et al.
Published: (2025)
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
by: Stollenwerk, Felix
Published: (2025)
by: Stollenwerk, Felix
Published: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
by: Kapadia, Shashank, et al.
Published: (2026)
by: Kapadia, Shashank, et al.
Published: (2026)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Thinking Deeper, Not Longer: Depth-Recurrent Transformers for Compositional Generalization
by: Chen, Hung-Hsuan
Published: (2026)
by: Chen, Hung-Hsuan
Published: (2026)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
by: Li, Jin, et al.
Published: (2025)
by: Li, Jin, et al.
Published: (2025)
Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Hidden Heroes and Gradient Bloats: Layer-Wise Redundancy Inverts Attribution in Transformers
by: Ye, Donald
Published: (2026)
by: Ye, Donald
Published: (2026)
CAST: Achieving Stable LLM-based Text Analysis for Data Analytics
by: Xie, Jinxiang, et al.
Published: (2026)
by: Xie, Jinxiang, et al.
Published: (2026)
Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuning
by: Zou, Jiaru, et al.
Published: (2025)
by: Zou, Jiaru, et al.
Published: (2025)
What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
by: Li, Ming, et al.
Published: (2024)
by: Li, Ming, et al.
Published: (2024)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
by: Gupta, Akshat, et al.
Published: (2024)
by: Gupta, Akshat, et al.
Published: (2024)
LoRAP: Transformer Sub-Layers Deserve Differentiated Structured Compression for Large Language Models
by: Li, Guangyan, et al.
Published: (2024)
by: Li, Guangyan, et al.
Published: (2024)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
RIPPLECOT: Amplifying Ripple Effect of Knowledge Editing in Language Models via Chain-of-Thought In-Context Learning
by: Zhao, Zihao, et al.
Published: (2024)
by: Zhao, Zihao, et al.
Published: (2024)
Revisiting the Shape Convention of Transformer Language Models
by: Liao, Feng-Ting, et al.
Published: (2026)
by: Liao, Feng-Ting, et al.
Published: (2026)
ALTA: Compiler-Based Analysis of Transformers
by: Shaw, Peter, et al.
Published: (2024)
by: Shaw, Peter, et al.
Published: (2024)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization
by: Bidusa, Or Raphael, et al.
Published: (2025)
by: Bidusa, Or Raphael, et al.
Published: (2025)
How Instruction and Reasoning Data shape Post-Training: Data Quality through the Lens of Layer-wise Gradients
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
by: Vincent, Sebastian, et al.
Published: (2023)
by: Vincent, Sebastian, et al.
Published: (2023)
Controlled Generation for Private Synthetic Text
by: Zhao, Zihao, et al.
Published: (2025)
by: Zhao, Zihao, et al.
Published: (2025)
OxonFair: A Flexible Toolkit for Algorithmic Fairness
by: Delaney, Eoin, et al.
Published: (2024)
by: Delaney, Eoin, et al.
Published: (2024)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
Understanding In-context Learning of Addition via Activation Subspaces
by: Hu, Xinyan, et al.
Published: (2025)
by: Hu, Xinyan, et al.
Published: (2025)
Similar Items
-
SCALPEL: Selective Capability Ablation via Low-rank Parameter Editing for Large Language Model Interpretability Analysis
by: Fu, Zihao, et al.
Published: (2026) -
Understanding Scaling Laws with Statistical and Approximation Theory for Transformer Neural Networks on Intrinsically Low-dimensional Data
by: Havrilla, Alex, et al.
Published: (2024) -
Understanding Transformers via N-gram Statistics
by: Nguyen, Timothy
Published: (2024) -
Compositional Generalization from Learned Skills via CoT Training: A Theoretical and Structural Analysis for Reasoning
by: Yao, Xinhao, et al.
Published: (2025) -
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)