Wonderful Matrices: Combining for a More Efficient and Effective Foundation Model Architecture
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Jingze, Wu, Bingheng |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Wonderful Matrices: More Efficient and Effective Architecture for Language Modeling Tasks
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025)
by: Shi, Jingze, et al.
Published: (2025)
TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding
by: Wu, Bingheng, et al.
Published: (2025)
by: Wu, Bingheng, et al.
Published: (2025)
Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
by: Song, Jingze, et al.
Published: (2026)
by: Song, Jingze, et al.
Published: (2026)
OTCE: Hybrid SSM and Attention with Cross Domain Mixture of Experts to construct Observer-Thinker-Conceiver-Expresser
by: Shi, Jingze, et al.
Published: (2024)
by: Shi, Jingze, et al.
Published: (2024)
When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Parameter-Efficient Fine-Tuning for Foundation Models
by: Zhang, Dan, et al.
Published: (2025)
by: Zhang, Dan, et al.
Published: (2025)
Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
by: Chen, Yingfa, et al.
Published: (2026)
by: Chen, Yingfa, et al.
Published: (2026)
Assessing the Portability of Parameter Matrices Trained by Parameter-Efficient Finetuning Methods
by: Sabry, Mohammed, et al.
Published: (2024)
by: Sabry, Mohammed, et al.
Published: (2024)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models
by: Noukhovitch, Michael, et al.
Published: (2024)
by: Noukhovitch, Michael, et al.
Published: (2024)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
by: Wu, Bohao, et al.
Published: (2025)
by: Wu, Bohao, et al.
Published: (2025)
ABBA-Adapters: Efficient and Expressive Fine-Tuning of Foundation Models
by: Singhal, Raghav, et al.
Published: (2025)
by: Singhal, Raghav, et al.
Published: (2025)
Towards More Effective Table-to-Text Generation: Assessing In-Context Learning and Self-Evaluation with Open-Source Models
by: Iravani, Sahar, et al.
Published: (2024)
by: Iravani, Sahar, et al.
Published: (2024)
Split, Unlearn, Merge: Leveraging Data Attributes for More Effective Unlearning in LLMs
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
by: Kadhe, Swanand Ravindra, et al.
Published: (2024)
Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search
by: Gu, Yuxian, et al.
Published: (2025)
by: Gu, Yuxian, et al.
Published: (2025)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
by: Wang, Xiao
Published: (2026)
by: Wang, Xiao
Published: (2026)
In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
by: Liu, Sheng, et al.
Published: (2023)
by: Liu, Sheng, et al.
Published: (2023)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
by: Ge, Albert, et al.
Published: (2025)
by: Ge, Albert, et al.
Published: (2025)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
by: Yue, Yuxuan, et al.
Published: (2024)
by: Yue, Yuxuan, et al.
Published: (2024)
More Bang for the Buck: Process Reward Modeling with Entropy-Driven Uncertainty
by: Cao, Lang, et al.
Published: (2025)
by: Cao, Lang, et al.
Published: (2025)
MIO: A Foundation Model on Multimodal Tokens
by: Wang, Zekun, et al.
Published: (2024)
by: Wang, Zekun, et al.
Published: (2024)
Tool Learning with Foundation Models
by: Qin, Yujia, et al.
Published: (2023)
by: Qin, Yujia, et al.
Published: (2023)
KGLens: Towards Efficient and Effective Knowledge Probing of Large Language Models with Knowledge Graphs
by: Zheng, Shangshang, et al.
Published: (2023)
by: Zheng, Shangshang, et al.
Published: (2023)
Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
by: He, Yufei, et al.
Published: (2025)
by: He, Yufei, et al.
Published: (2025)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
by: Zhang, Yifei, et al.
Published: (2024)
by: Zhang, Yifei, et al.
Published: (2024)
Effectively Controlling Reasoning Models through Thinking Intervention
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Foundations of Large Language Models
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
Apple Intelligence Foundation Language Models
by: Gunter, Tom, et al.
Published: (2024)
by: Gunter, Tom, et al.
Published: (2024)
Weaver: Foundation Models for Creative Writing
by: Wang, Tiannan, et al.
Published: (2024)
by: Wang, Tiannan, et al.
Published: (2024)
COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees
by: Wang, Zhiyuan, et al.
Published: (2025)
by: Wang, Zhiyuan, et al.
Published: (2025)
Elastic Architecture Search for Efficient Language Models
by: Wang, Shang
Published: (2025)
by: Wang, Shang
Published: (2025)
Less is More: Local Intrinsic Dimensions of Contextual Language Models
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
by: Ruppik, Benjamin Matthias, et al.
Published: (2025)
EBFT: Effective and Block-Wise Fine-Tuning for Sparse LLMs
by: Guo, Song, et al.
Published: (2024)
by: Guo, Song, et al.
Published: (2024)
Heterogeneous Scientific Foundation Model Collaboration
by: Li, Zihao, et al.
Published: (2026)
by: Li, Zihao, et al.
Published: (2026)
Every Attention Matters: An Efficient Hybrid Architecture for Long-Context Reasoning
by: Ling Team, et al.
Published: (2025)
by: Ling Team, et al.
Published: (2025)
When More is Less: Understanding Chain-of-Thought Length in LLMs
by: Wu, Yuyang, et al.
Published: (2025)
by: Wu, Yuyang, et al.
Published: (2025)
LoRA-Mini : Adaptation Matrices Decomposition and Selective Training
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
by: Yang, Yazheng, et al.
Published: (2023)
by: Yang, Yazheng, et al.
Published: (2023)
Similar Items
-
Wonderful Matrices: More Efficient and Effective Architecture for Language Modeling Tasks
by: Shi, Jingze, et al.
Published: (2024) -
Trainable Dynamic Mask Sparse Attention
by: Shi, Jingze, et al.
Published: (2025) -
TransXSSM: A Hybrid Transformer State Space Model with Unified Rotary Position Embedding
by: Wu, Bingheng, et al.
Published: (2025) -
Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxies
by: Song, Jingze, et al.
Published: (2026) -
OTCE: Hybrid SSM and Attention with Cross Domain Mixture of Experts to construct Observer-Thinker-Conceiver-Expresser
by: Shi, Jingze, et al.
Published: (2024)