TeamFormer: Shallow Parallel Transformers with Progressive Approximation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Wei, Wei, Xiao-Yong, Li, Qing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Approximation in Deep Residual Networks: Theory and Validation
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Dynamic Universal Approximation Theory: The Basic Theory for Transformer-based Large Language Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
by: Cutler, Dylan, et al.
Published: (2025)
by: Cutler, Dylan, et al.
Published: (2025)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
PolyFormer: Scalable Node-wise Filters via Polynomial Graph Transformer
by: Ma, Jiahong, et al.
Published: (2024)
by: Ma, Jiahong, et al.
Published: (2024)
ImputeFormer: Low Rankness-Induced Transformers for Generalizable Spatiotemporal Imputation
by: Nie, Tong, et al.
Published: (2023)
by: Nie, Tong, et al.
Published: (2023)
FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification
by: Liu, Liming, et al.
Published: (2025)
by: Liu, Liming, et al.
Published: (2025)
Mol-Debate: Multi-Agent Debate Improves Structural Reasoning in Molecular Design
by: Zhang, Wengyu, et al.
Published: (2026)
by: Zhang, Wengyu, et al.
Published: (2026)
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
VT-Former: Efffcient Transformer-based Decoder for Varshamov-Tenengolts Codes
by: Wei, Yali, et al.
Published: (2025)
by: Wei, Yali, et al.
Published: (2025)
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
by: Li, Zhonghao, et al.
Published: (2025)
by: Li, Zhonghao, et al.
Published: (2025)
E2Former: An Efficient and Equivariant Transformer with Linear-Scaling Tensor Products
by: Li, Yunyang, et al.
Published: (2025)
by: Li, Yunyang, et al.
Published: (2025)
IceFormer: Accelerated Inference with Long-Sequence Transformers on CPUs
by: Mao, Yuzhen, et al.
Published: (2024)
by: Mao, Yuzhen, et al.
Published: (2024)
TokenFormer: Rethinking Transformer Scaling with Tokenized Model Parameters
by: Wang, Haiyang, et al.
Published: (2024)
by: Wang, Haiyang, et al.
Published: (2024)
CausalFormer: An Interpretable Transformer for Temporal Causal Discovery
by: Kong, Lingbai, et al.
Published: (2024)
by: Kong, Lingbai, et al.
Published: (2024)
CoFormer: Collaborating with Heterogeneous Edge Devices for Scalable Transformer Inference
by: Xu, Guanyu, et al.
Published: (2025)
by: Xu, Guanyu, et al.
Published: (2025)
TouchFormer: A Robust Transformer-based Framework for Multimodal Material Perception
by: Lyu, Kailin, et al.
Published: (2025)
by: Lyu, Kailin, et al.
Published: (2025)
TinyFusion: Diffusion Transformers Learned Shallow
by: Fang, Gongfan, et al.
Published: (2024)
by: Fang, Gongfan, et al.
Published: (2024)
RingFormer: A Neural Vocoder with Ring Attention and Convolution-Augmented Transformer
by: Hong, Seongho, et al.
Published: (2025)
by: Hong, Seongho, et al.
Published: (2025)
Statistical Guarantees for Approximate Stationary Points of Shallow Neural Networks
by: Taheri, Mahsa, et al.
Published: (2022)
by: Taheri, Mahsa, et al.
Published: (2022)
Why Shallow Networks Struggle to Approximate and Learn High Frequencies
by: Zhang, Shijun, et al.
Published: (2023)
by: Zhang, Shijun, et al.
Published: (2023)
WaveFormer: Wavelet Embedding Transformer for Biomedical Signals
by: Irani, Habib, et al.
Published: (2026)
by: Irani, Habib, et al.
Published: (2026)
Chain-of-Thought Enhanced Shallow Transformers for Wireless Symbol Detection
by: Fan, Li, et al.
Published: (2025)
by: Fan, Li, et al.
Published: (2025)
Parallel Layer Normalization for Universal Approximation
by: Ni, Yunhao, et al.
Published: (2025)
by: Ni, Yunhao, et al.
Published: (2025)
Approximate Top-$k$ for Increased Parallelism
by: Key, Oscar, et al.
Published: (2024)
by: Key, Oscar, et al.
Published: (2024)
Elastic Multi-Gradient Descent for Parallel Continual Learning
by: Lyu, Fan, et al.
Published: (2024)
by: Lyu, Fan, et al.
Published: (2024)
ContiFormer: Continuous-Time Transformer for Irregular Time Series Modeling
by: Chen, Yuqi, et al.
Published: (2024)
by: Chen, Yuqi, et al.
Published: (2024)
TyphoFormer: Language-Augmented Transformer for Accurate Typhoon Track Forecasting
by: Li, Lincan, et al.
Published: (2025)
by: Li, Lincan, et al.
Published: (2025)
TimeFormer: Transformer with Attention Modulation Empowered by Temporal Characteristics for Time Series Forecasting
by: Liu, Zhipeng, et al.
Published: (2025)
by: Liu, Zhipeng, et al.
Published: (2025)
Parallel BiLSTM-Transformer networks for forecasting chaotic dynamics
by: Ma, Junwen, et al.
Published: (2025)
by: Ma, Junwen, et al.
Published: (2025)
Transformers Meet In-Context Learning: A Universal Approximation Theory
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
TabTreeFormer: Tabular Data Generation Using Hybrid Tree-Transformer
by: Li, Jiayu, et al.
Published: (2025)
by: Li, Jiayu, et al.
Published: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
by: Zhou, Jingbo, et al.
Published: (2026)
by: Zhou, Jingbo, et al.
Published: (2026)
Massively Parallel Expectation Maximization For Approximate Posteriors
by: Heap, Thomas, et al.
Published: (2025)
by: Heap, Thomas, et al.
Published: (2025)
IsingFormer: Augmenting Parallel Tempering With Learned Proposals
by: Bunaiyan, Saleh, et al.
Published: (2025)
by: Bunaiyan, Saleh, et al.
Published: (2025)
NN-Former: Rethinking Graph Structure in Neural Architecture Representation
by: Xu, Ruihan, et al.
Published: (2025)
by: Xu, Ruihan, et al.
Published: (2025)
CopRA: A Progressive LoRA Training Strategy
by: Zhuang, Zhan, et al.
Published: (2024)
by: Zhuang, Zhan, et al.
Published: (2024)
AlgoFormer: An Efficient Transformer Framework with Algorithmic Structures
by: Gao, Yihang, et al.
Published: (2024)
by: Gao, Yihang, et al.
Published: (2024)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
by: Zou, Shihao, et al.
Published: (2025)
by: Zou, Shihao, et al.
Published: (2025)
Similar Items
-
Progressive Approximation in Deep Residual Networks: Theory and Validation
by: Wang, Wei, et al.
Published: (2026) -
Dynamic Universal Approximation Theory: Foundations for Parallelism in Neural Networks
by: Wang, Wei, et al.
Published: (2024) -
Dynamic Universal Approximation Theory: The Basic Theory for Transformer-based Large Language Models
by: Wang, Wei, et al.
Published: (2024) -
StagFormer: Time Staggering Transformer Decoding for RunningLayers In Parallel
by: Cutler, Dylan, et al.
Published: (2025) -
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)