How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
Fuente:
arXiv
Saved in:
| Main Authors: | Lu, Xin, Zhao, Yanyan, Wei, Si, Wang, Shijin, Qin, Bing, Liu, Ting |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
by: Lu, Xin, et al.
Published: (2024)
by: Lu, Xin, et al.
Published: (2024)
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
by: Yi, Rongjie, et al.
Published: (2024)
by: Yi, Rongjie, et al.
Published: (2024)
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
by: Tang, Tianyi, et al.
Published: (2024)
by: Tang, Tianyi, et al.
Published: (2024)
Design Principle Transfer in Neural Architecture Search via Large Language Models
by: Zhou, Xun, et al.
Published: (2024)
by: Zhou, Xun, et al.
Published: (2024)
The Lazuli Space Observatory: Architecture & Capabilities
by: Roy, Arpita, et al.
Published: (2026)
by: Roy, Arpita, et al.
Published: (2026)
IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities
by: Wang, Bin, et al.
Published: (2024)
by: Wang, Bin, et al.
Published: (2024)
Exploring the Capabilities of Large Language Models for Generating Diverse Design Solutions
by: Ma, Kevin, et al.
Published: (2024)
by: Ma, Kevin, et al.
Published: (2024)
AfriqueLLM: How Data Mixing and Model Architecture Impact Continued Pre-training for African Languages
by: Yu, Hao, et al.
Published: (2026)
by: Yu, Hao, et al.
Published: (2026)
Companion Materials for: Epistemological Sequencing as an Architectural Principle for Orchestrating Reasoning in Language Models
by: Manucci, Marcelo
Published: (2026)
by: Manucci, Marcelo
Published: (2026)
Exploring the Adversarial Capabilities of Large Language Models
by: Struppek, Lukas, et al.
Published: (2024)
by: Struppek, Lukas, et al.
Published: (2024)
Sequence-to-Sequence Spanish Pre-trained Language Models
by: Araujo, Vladimir, et al.
Published: (2023)
by: Araujo, Vladimir, et al.
Published: (2023)
From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
by: Huang, Lanxiao, et al.
Published: (2025)
by: Huang, Lanxiao, et al.
Published: (2025)
Model‐Based Modularization Approach for Reconfigurable System Architecture Design
by: Zhemei Fang, et al.
Published: (2026)
by: Zhemei Fang, et al.
Published: (2026)
Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models
by: Wu, Di, et al.
Published: (2024)
by: Wu, Di, et al.
Published: (2024)
A Layered Architecture for Developing and Enhancing Capabilities in Large Language Model-based Software Systems
by: Zhang, Dawen, et al.
Published: (2024)
by: Zhang, Dawen, et al.
Published: (2024)
Small Language Models for Agentic Systems: A Survey of Architectures, Capabilities, and Deployment Trade offs
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Structural Pruning of Pre-trained Language Models via Neural Architecture Search
by: Klein, Aaron, et al.
Published: (2024)
by: Klein, Aaron, et al.
Published: (2024)
AEROS: A Single-Agent Operating Architecture with Embodied Capability Modules
by: Qin, Xue, et al.
Published: (2026)
by: Qin, Xue, et al.
Published: (2026)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
by: Bhattamishra, Satwik, et al.
Published: (2024)
by: Bhattamishra, Satwik, et al.
Published: (2024)
An AI Architecture with the Capability to Explain Recognition Results
by: Whitten, Paul, et al.
Published: (2024)
by: Whitten, Paul, et al.
Published: (2024)
CARE-Bench: A Benchmark of Diverse Client Simulations Guided by Expert Principles for Evaluating LLMs in Psychological Counseling
by: Wang, Bichen, et al.
Published: (2025)
by: Wang, Bichen, et al.
Published: (2025)
How Does Cognitive Capability and Personality Influence Problem Solving in Coding Interview Puzzles?
by: Hidellaarachchi, Dulaji, et al.
Published: (2025)
by: Hidellaarachchi, Dulaji, et al.
Published: (2025)
Exploring State Tracking Capabilities of Large Language Models
by: Rezaee, Kiamehr, et al.
Published: (2025)
by: Rezaee, Kiamehr, et al.
Published: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Exploring the Learning Capabilities of Language Models using LEVERWORLDS
by: Wagner, Eitan, et al.
Published: (2024)
by: Wagner, Eitan, et al.
Published: (2024)
Retrievit: In-context Retrieval Capabilities of Transformers, State Space Models, and Hybrid Architectures
by: Pantazopoulos, Georgios, et al.
Published: (2026)
by: Pantazopoulos, Georgios, et al.
Published: (2026)
Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models
by: Shi, Chengshuai, et al.
Published: (2024)
by: Shi, Chengshuai, et al.
Published: (2024)
Designing Intelligent Enterprise Agents: A Capability-Aligned Multi-Agent Architecture
by: deVadoss, John
Published: (2026)
by: deVadoss, John
Published: (2026)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
Probing Audio-Generation Capabilities of Text-Based Language Models
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
by: Anbazhagan, Arjun Prasaath, et al.
Published: (2025)
Hephaestus: Improving Fundamental Agent Capabilities of Large Language Models through Continual Pre-Training
by: Zhuang, Yuchen, et al.
Published: (2025)
by: Zhuang, Yuchen, et al.
Published: (2025)
Improving the Downstream Performance of Mixture-of-Experts Transformers via Weak Vanilla Transformers
by: Lu, Xin, et al.
Published: (2024)
by: Lu, Xin, et al.
Published: (2024)
Assessing Capability Complexity Using Enterprise Architecture Framework
by: Javad Bakhshi, et al.
Published: (2026)
by: Javad Bakhshi, et al.
Published: (2026)
An AI Architecture with the Capability to Classify and Explain Hardware Trojans
by: Whitten, Paul, et al.
Published: (2024)
by: Whitten, Paul, et al.
Published: (2024)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
Probing Language Models for Pre-training Data Detection
by: Liu, Zhenhua, et al.
Published: (2024)
by: Liu, Zhenhua, et al.
Published: (2024)
Generative Reliability-Based Design Optimization Using In-Context Learning Capabilities of Large Language Models
by: Jiang, Zhonglin, et al.
Published: (2025)
by: Jiang, Zhonglin, et al.
Published: (2025)
Exploring the Privacy Protection Capabilities of Chinese Large Language Models
by: Yang, Yuqi, et al.
Published: (2024)
by: Yang, Yuqi, et al.
Published: (2024)
Exploring the Word Sense Disambiguation Capabilities of Large Language Models
by: Basile, Pierpaolo, et al.
Published: (2025)
by: Basile, Pierpaolo, et al.
Published: (2025)
The Foundational Capabilities of Large Language Models in Predicting Postoperative Risks Using Clinical Notes
by: Alba, Charles, et al.
Published: (2024)
by: Alba, Charles, et al.
Published: (2024)
Similar Items
-
How does Architecture Influence the Base Capabilities of Pre-trained Language Models? A Case Study Based on FFN-Wider and MoE Transformers
by: Lu, Xin, et al.
Published: (2024) -
PhoneLM:an Efficient and Capable Small Language Model Family through Principled Pre-training
by: Yi, Rongjie, et al.
Published: (2024) -
Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models
by: Tang, Tianyi, et al.
Published: (2024) -
Design Principle Transfer in Neural Architecture Search via Large Language Models
by: Zhou, Xun, et al.
Published: (2024) -
The Lazuli Space Observatory: Architecture & Capabilities
by: Roy, Arpita, et al.
Published: (2026)