Geometric Interpretation of Layer Normalization and a Comparative Analysis with RMSNorm
Fuente:
arXiv
Salvato in:
| Autori principali: | Gupta, Akshat, Ozdemir, Atahan, Anumanchipalli, Gopala |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Unified Framework for Model Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
di: Yoon, Junsang, et al.
Pubblicazione: (2024)
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
di: Abdi, Immanuel, et al.
Pubblicazione: (2026)
di: Abdi, Immanuel, et al.
Pubblicazione: (2026)
Lifelong Knowledge Editing requires Better Regularization
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
Rebuilding ROME : Resolving Model Collapse during Sequential Model Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
Self-Assessment Tests are Unreliable Measures of LLM Personality
di: Gupta, Akshat, et al.
Pubblicazione: (2023)
di: Gupta, Akshat, et al.
Pubblicazione: (2023)
FineZip : Pushing the Limits of Large Language Models for Practical Lossless Text Compression
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
di: Mittu, Fazal, et al.
Pubblicazione: (2024)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
di: Gupta, Akshat, et al.
Pubblicazione: (2024)
How Do LLMs Use Their Depth?
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
Efficient Knowledge Editing via Minimal Precomputation
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
di: Gupta, Akshat, et al.
Pubblicazione: (2025)
The Oracle Has Spoken: A Multi-Aspect Evaluation of Dialogue in Pythia
di: Chen, Zixun, et al.
Pubblicazione: (2025)
di: Chen, Zixun, et al.
Pubblicazione: (2025)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
A Persona-Based Evaluation Framework for Pluralistic Alignment in Generative AI
di: Karagoz, Atahan
Pubblicazione: (2026)
di: Karagoz, Atahan
Pubblicazione: (2026)
Unsupervised TTS Acoustic Modeling for TTS with Conditional Disentangled Sequential VAE
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
di: Lian, Jiachen, et al.
Pubblicazione: (2022)
PokerBench: Training Large Language Models to become Professional Poker Players
di: Zhuang, Richard, et al.
Pubblicazione: (2025)
di: Zhuang, Richard, et al.
Pubblicazione: (2025)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
di: Kim, Jeonghoon, et al.
Pubblicazione: (2025)
di: Kim, Jeonghoon, et al.
Pubblicazione: (2025)
Identifying Multiple Personalities in Large Language Models with External Evaluation
di: Song, Xiaoyang, et al.
Pubblicazione: (2024)
di: Song, Xiaoyang, et al.
Pubblicazione: (2024)
On the Mathematical Relationship Between Layer Normalization and Dynamic Activation Functions
di: Stollenwerk, Felix
Pubblicazione: (2025)
di: Stollenwerk, Felix
Pubblicazione: (2025)
On the Geometric Structure of Layer Updates in Deep Language Models
di: Yoo, Jun-Sik
Pubblicazione: (2026)
di: Yoo, Jun-Sik
Pubblicazione: (2026)
Cross-Layer Discrete Concept Discovery for Interpreting Language Models
di: Garg, Ankur, et al.
Pubblicazione: (2025)
di: Garg, Ankur, et al.
Pubblicazione: (2025)
Concept Layers: Enhancing Interpretability and Intervenability via LLM Conceptualization
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
di: Bidusa, Or Raphael, et al.
Pubblicazione: (2025)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
di: Askari, Hadi, et al.
Pubblicazione: (2025)
di: Askari, Hadi, et al.
Pubblicazione: (2025)
Comparing Feature Importance and Rule Extraction for Interpretability on Text Data
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
di: Lopardo, Gianluigi, et al.
Pubblicazione: (2022)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
di: Sim, Woo Seob, et al.
Pubblicazione: (2026)
di: Sim, Woo Seob, et al.
Pubblicazione: (2026)
LayerSkip: Enabling Early Exit Inference and Self-Speculative Decoding
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2024)
di: Elhoushi, Mostafa, et al.
Pubblicazione: (2024)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
di: Méloux, Maxime, et al.
Pubblicazione: (2025)
Ethics and Technical Aspects of Generative AI Models in Digital Content Creation
di: Karagoz, Atahan
Pubblicazione: (2024)
di: Karagoz, Atahan
Pubblicazione: (2024)
A Comparative Study of Neurosymbolic AI Approaches to Interpretable Logical Reasoning
di: Chen, Michael K.
Pubblicazione: (2025)
di: Chen, Michael K.
Pubblicazione: (2025)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
di: Joshi, Abhinav, et al.
Pubblicazione: (2024)
Limitations of Normalization in Attention Mechanism
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
di: Mudarisov, Timur, et al.
Pubblicazione: (2025)
MIB: A Mechanistic Interpretability Benchmark
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
di: Shu, Dong, et al.
Pubblicazione: (2024)
di: Shu, Dong, et al.
Pubblicazione: (2024)
Geometric-disentangelment Unlearning
di: Zhou, Duo, et al.
Pubblicazione: (2025)
di: Zhou, Duo, et al.
Pubblicazione: (2025)
CAST: Compositional Analysis via Spectral Tracking for Understanding Transformer Layer Functions
di: Fu, Zihao, et al.
Pubblicazione: (2025)
di: Fu, Zihao, et al.
Pubblicazione: (2025)
Optimizing Large Language Models for Turkish: New Methodologies in Corpus Selection and Training
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
di: Kesgin, H. Toprak, et al.
Pubblicazione: (2024)
Layer by Layer: Uncovering Hidden Representations in Language Models
di: Skean, Oscar, et al.
Pubblicazione: (2025)
di: Skean, Oscar, et al.
Pubblicazione: (2025)
The Effectiveness of LLMs as Annotators: A Comparative Overview and Empirical Analysis of Direct Representation
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
di: Pavlovic, Maja, et al.
Pubblicazione: (2024)
Discourse-Aware In-Context Learning for Temporal Expression Normalization
di: Gautam, Akash Kumar, et al.
Pubblicazione: (2024)
di: Gautam, Akash Kumar, et al.
Pubblicazione: (2024)
Toward a Functional Geometric Algebra for Natural Language Semantics
di: Pustejovsky, James
Pubblicazione: (2026)
di: Pustejovsky, James
Pubblicazione: (2026)
Documenti analoghi
-
A Unified Framework for Model Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2024) -
Is Bigger Edit Batch Size Always Better? -- An Empirical Study on Model Editing with Llama-3
di: Yoon, Junsang, et al.
Pubblicazione: (2024) -
Norm Growth and Stability Challenges in Localized Sequential Knowledge Editing
di: Gupta, Akshat, et al.
Pubblicazione: (2025) -
Evolutionary Strategies lead to Catastrophic Forgetting in LLMs
di: Abdi, Immanuel, et al.
Pubblicazione: (2026) -
Lifelong Knowledge Editing requires Better Regularization
di: Gupta, Akshat, et al.
Pubblicazione: (2025)