Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Allen-Zhu, Zeyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 3.2, Knowledge Manipulation
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023)
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024)
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
von: Ye, Tian, et al.
Veröffentlicht: (2024)
von: Ye, Tian, et al.
Veröffentlicht: (2024)
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems
von: Ye, Tian, et al.
Veröffentlicht: (2024)
von: Ye, Tian, et al.
Veröffentlicht: (2024)
Reverse Training to Nurse the Reversal Curse
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
von: Golovneva, Olga, et al.
Veröffentlicht: (2024)
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
von: Bae, Sangmin, et al.
Veröffentlicht: (2025)
von: Bae, Sangmin, et al.
Veröffentlicht: (2025)
Language Models over Canonical Byte-Pair Encodings
von: Vieira, Tim, et al.
Veröffentlicht: (2025)
von: Vieira, Tim, et al.
Veröffentlicht: (2025)
Architectural Foundations for the Large Language Model Infrastructures
von: Zhu, Hongyin
Veröffentlicht: (2024)
von: Zhu, Hongyin
Veröffentlicht: (2024)
Magical: Medical Lay Language Generation via Semantic Invariance and Layperson-tailored Adaptation
von: Liao, Weibin, et al.
Veröffentlicht: (2025)
von: Liao, Weibin, et al.
Veröffentlicht: (2025)
Model Editing with Canonical Examples
von: Hewitt, John, et al.
Veröffentlicht: (2024)
von: Hewitt, John, et al.
Veröffentlicht: (2024)
Beyond Magic Words: Sharpness-Aware Prompt Evolving for Robust Large Language Models with TARE
von: Wan, Guancheng, et al.
Veröffentlicht: (2025)
von: Wan, Guancheng, et al.
Veröffentlicht: (2025)
Similarity-Distance-Magnitude Language Models
von: Schmaltz, Allen
Veröffentlicht: (2025)
von: Schmaltz, Allen
Veröffentlicht: (2025)
Broken Tokens? Your Language Model can Secretly Handle Non-Canonical Tokenizations
von: Zheng, Brian Siyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Brian Siyuan, et al.
Veröffentlicht: (2025)
Structured In-context Environment Scaling for Large Language Model Reasoning
von: Yu, Peng, et al.
Veröffentlicht: (2025)
von: Yu, Peng, et al.
Veröffentlicht: (2025)
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
von: Lin, Zizhuo, et al.
Veröffentlicht: (2026)
von: Lin, Zizhuo, et al.
Veröffentlicht: (2026)
How Far Are Vision-Language Models from Constructing the Real World? A Benchmark for Physical Generative Reasoning
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
von: Yang, Luyu, et al.
Veröffentlicht: (2026)
Persona-Conditioned Risk Behavior in Large Language Models: A Simulated Gambling Study with GPT-4.1
von: Dubedy, Sankalp
Veröffentlicht: (2026)
von: Dubedy, Sankalp
Veröffentlicht: (2026)
Behind the Magic, MERLIM: Multi-modal Evaluation Benchmark for Large Image-Language Models
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
von: Villa, Andrés, et al.
Veröffentlicht: (2023)
Balancing Rigor and Utility: Mitigating Cognitive Biases in Large Language Models for Multiple-Choice Questions
von: Zhong, Hanyang, et al.
Veröffentlicht: (2024)
von: Zhong, Hanyang, et al.
Veröffentlicht: (2024)
Language-Specific Layer Matters: Efficient Multilingual Enhancement for Large Vision-Language Models
von: Fan, Yuchun, et al.
Veröffentlicht: (2025)
von: Fan, Yuchun, et al.
Veröffentlicht: (2025)
Nanbeige4.1-3B: A Small General Model that Reasons, Aligns, and Acts
von: Yang, Chen, et al.
Veröffentlicht: (2026)
von: Yang, Chen, et al.
Veröffentlicht: (2026)
Design Principle Transfer in Neural Architecture Search via Large Language Models
von: Zhou, Xun, et al.
Veröffentlicht: (2024)
von: Zhou, Xun, et al.
Veröffentlicht: (2024)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025)
von: Lu, Xin, et al.
Veröffentlicht: (2025)
You Only Cache Once: Decoder-Decoder Architectures for Language Models
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
The Compressor-Retriever Architecture for Language Model OS
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
von: Yang, Yuan, et al.
Veröffentlicht: (2024)
Jailbreaking LLMs' Safeguard with Universal Magic Words for Text Embedding Models
von: Liang, Haoyu, et al.
Veröffentlicht: (2025)
von: Liang, Haoyu, et al.
Veröffentlicht: (2025)
Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
von: Zhu, Jinchang, et al.
Veröffentlicht: (2026)
A Reference Architecture for Designing Foundation Model based Systems
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
von: Lu, Qinghua, et al.
Veröffentlicht: (2023)
NL2GenSym: Natural Language to Generative Symbolic Rules for SOAR Cognitive Architecture via Large Language Models
von: Yuan, Fang, et al.
Veröffentlicht: (2025)
von: Yuan, Fang, et al.
Veröffentlicht: (2025)
FastSpell: the LangId Magic Spell
von: Bañón, Marta, et al.
Veröffentlicht: (2024)
von: Bañón, Marta, et al.
Veröffentlicht: (2024)
Curriculum-Guided Layer Scaling for Language Model Pretraining
von: Singh, Karanpartap, et al.
Veröffentlicht: (2025)
von: Singh, Karanpartap, et al.
Veröffentlicht: (2025)
Architecture, Not Scale: Circuit Localization in Large Language Models
von: Venkatesh, Sohan
Veröffentlicht: (2026)
von: Venkatesh, Sohan
Veröffentlicht: (2026)
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
von: Li, You, et al.
Veröffentlicht: (2025)
von: Li, You, et al.
Veröffentlicht: (2025)
Multilingual Contrastive Decoding via Language-Agnostic Layers Skipping
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
von: Zhu, Wenhao, et al.
Veröffentlicht: (2024)
Scaling Embedding Layers in Language Models
von: Yu, Da, et al.
Veröffentlicht: (2025)
von: Yu, Da, et al.
Veröffentlicht: (2025)
Foundations of Large Language Model Compression -- Part 1: Weight Quantization
von: Young, Sean I.
Veröffentlicht: (2024)
von: Young, Sean I.
Veröffentlicht: (2024)
Magic Markup: Maintaining Document-External Markup with an LLM
von: Misback, Edward, et al.
Veröffentlicht: (2024)
von: Misback, Edward, et al.
Veröffentlicht: (2024)
MCP4IFC: IFC-Based Building Design Using Large Language Models
von: Nithyanantham, Bharathi Kannan, et al.
Veröffentlicht: (2025)
von: Nithyanantham, Bharathi Kannan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Physics of Language Models: Part 1, Learning Hierarchical Language Structures
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023) -
Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023) -
Physics of Language Models: Part 3.2, Knowledge Manipulation
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2023) -
Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
von: Allen-Zhu, Zeyuan, et al.
Veröffentlicht: (2024) -
Physics of Language Models: Part 2.1, Grade-School Math and the Hidden Reasoning Process
von: Ye, Tian, et al.
Veröffentlicht: (2024)