Transformer Block Coupling and its Correlation with Generalization in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Aubry, Murdock, Meng, Haoming, Sugolov, Anton, Papyan, Vardan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
von: Zhang, Stephen, et al.
Veröffentlicht: (2025)
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
von: Zhang, Stephen, et al.
Veröffentlicht: (2024)
Pushing Boundaries: Mixup's Influence on Neural Collapse
von: Fisher, Quinn, et al.
Veröffentlicht: (2024)
von: Fisher, Quinn, et al.
Veröffentlicht: (2024)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
Use Your INSTINCT: INSTruction optimization for LLMs usIng Neural bandits Coupled with Transformers
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
von: Lin, Xiaoqiang, et al.
Veröffentlicht: (2023)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
von: Guo, Song, et al.
Veröffentlicht: (2024)
von: Guo, Song, et al.
Veröffentlicht: (2024)
MoBA: Mixture of Block Attention for Long-Context LLMs
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
von: Lu, Enzhe, et al.
Veröffentlicht: (2025)
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
von: Kowshik, Suhas S, et al.
Veröffentlicht: (2024)
von: Kowshik, Suhas S, et al.
Veröffentlicht: (2024)
BlockCert: Certified Blockwise Extraction of Transformer Mechanisms
von: Andric, Sandro
Veröffentlicht: (2025)
von: Andric, Sandro
Veröffentlicht: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
von: Wu, Robert, et al.
Veröffentlicht: (2024)
von: Wu, Robert, et al.
Veröffentlicht: (2024)
Position Coupling: Improving Length Generalization of Arithmetic Transformers Using Task Structure
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
von: Cho, Hanseul, et al.
Veröffentlicht: (2024)
Ensemble Learning for Large Language Models in Text and Code Generation: A Survey
von: Ashiga, Mari, et al.
Veröffentlicht: (2025)
von: Ashiga, Mari, et al.
Veröffentlicht: (2025)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
von: Ho, Namgyu, et al.
Veröffentlicht: (2024)
PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
von: Li, Haoming, et al.
Veröffentlicht: (2025)
von: Li, Haoming, et al.
Veröffentlicht: (2025)
ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
von: Shopkhoev, Dmitriy, et al.
Veröffentlicht: (2025)
Laplacian Heads Improve Transformers by Smoothing Token Representations
von: Zhang, Yuchong, et al.
Veröffentlicht: (2026)
von: Zhang, Yuchong, et al.
Veröffentlicht: (2026)
Transformer-Squared: Self-adaptive LLMs
von: Sun, Qi, et al.
Veröffentlicht: (2025)
von: Sun, Qi, et al.
Veröffentlicht: (2025)
Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
von: Goldman, Omer, et al.
Veröffentlicht: (2024)
Evaluating the Generalization Ability of Quantized LLMs: Benchmark, Analysis, and Toolbox
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
von: Liu, Yijun, et al.
Veröffentlicht: (2024)
LASP: Surveying the State-of-the-Art in Large Language Model-Assisted AI Planning
von: Li, Haoming, et al.
Veröffentlicht: (2024)
von: Li, Haoming, et al.
Veröffentlicht: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
von: Chen, Junqi, et al.
Veröffentlicht: (2026)
When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
von: Li, Jin, et al.
Veröffentlicht: (2025)
von: Li, Jin, et al.
Veröffentlicht: (2025)
Your Transformer is Secretly Linear
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
von: Razzhigaev, Anton, et al.
Veröffentlicht: (2024)
BgGPT 1.0: Extending English-centric LLMs to other languages
von: Alexandrov, Anton, et al.
Veröffentlicht: (2024)
von: Alexandrov, Anton, et al.
Veröffentlicht: (2024)
Block-Attention for Efficient Prefilling
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
von: Ma, Dongyang, et al.
Veröffentlicht: (2024)
Evaluation of Large Language Models via Coupled Token Generation
von: Benz, Nina Corvelo, et al.
Veröffentlicht: (2025)
von: Benz, Nina Corvelo, et al.
Veröffentlicht: (2025)
The Role of Sparsity for Length Generalization in Transformers
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
von: Golowich, Noah, et al.
Veröffentlicht: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
von: Andriushchenko, Maksym, et al.
Veröffentlicht: (2024)
Can LLMs Speak For Diverse People? Tuning LLMs via Debate to Generate Controllable Controversial Statements
von: Li, Ming, et al.
Veröffentlicht: (2024)
von: Li, Ming, et al.
Veröffentlicht: (2024)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
von: Reddy, Avinash, et al.
Veröffentlicht: (2026)
von: Reddy, Avinash, et al.
Veröffentlicht: (2026)
Trust the Batch, On- or Off-Policy: Adaptive Policy Optimization for RL Post-Training
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
von: Fakoor, Rasool, et al.
Veröffentlicht: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
Direct Behavior Optimization: Unlocking the Potential of Lightweight LLMs
von: Yang, Hongming, et al.
Veröffentlicht: (2025)
von: Yang, Hongming, et al.
Veröffentlicht: (2025)
Text-to-Stage: Spatial Layouts from Long-form Narratives
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2026)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2026)
Not All Layers of LLMs Are Necessary During Inference
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
von: Fan, Siqi, et al.
Veröffentlicht: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
von: Zhou, Yongchao, et al.
Veröffentlicht: (2024)
von: Zhou, Yongchao, et al.
Veröffentlicht: (2024)
Revisiting the Shape Convention of Transformer Language Models
von: Liao, Feng-Ting, et al.
Veröffentlicht: (2026)
von: Liao, Feng-Ting, et al.
Veröffentlicht: (2026)
Online Personalizing White-box LLMs Generation with Neural Bandits
von: Chen, Zekai, et al.
Veröffentlicht: (2024)
von: Chen, Zekai, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Attention Sinks: A 'Catch, Tag, Release' Mechanism for Embeddings
von: Zhang, Stephen, et al.
Veröffentlicht: (2025) -
OATS: Outlier-Aware Pruning Through Sparse and Low Rank Decomposition
von: Zhang, Stephen, et al.
Veröffentlicht: (2024) -
Sparsest Models Elude Pruning: An Exposé of Pruning's Current Capabilities
von: Zhang, Stephen, et al.
Veröffentlicht: (2024) -
Pushing Boundaries: Mixup's Influence on Neural Collapse
von: Fisher, Quinn, et al.
Veröffentlicht: (2024) -
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)