ComplexFormer: Disruptively Advancing Transformer Inference Ability via Head-Specific Complex Vector Attention
Fuente:
arXiv
Guardado en:
| Autores principales: | Shao, Jintian, Huang, Hongyi, Wu, Jiayi, Zhang, Beiwen, Wu, ZhiYu, Shan, You, Zheng, MingKai |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
EulerFormer: Sequential User Behavior Modeling with Complex Vector Attention
por: Tian, Zhen, et al.
Publicado: (2024)
por: Tian, Zhen, et al.
Publicado: (2024)
CipherFormer: Efficient Transformer Private Inference with Low Round Complexity
por: Wang, Weize, et al.
Publicado: (2024)
por: Wang, Weize, et al.
Publicado: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)
por: Yan, Ruiqing, et al.
Publicado: (2024)
Interactive Multi-Head Self-Attention with Linear Complexity
por: Kang, Hankyul, et al.
Publicado: (2024)
por: Kang, Hankyul, et al.
Publicado: (2024)
DUFOMap: Efficient Dynamic Awareness Mapping
por: Duberg, Daniel, et al.
Publicado: (2024)
por: Duberg, Daniel, et al.
Publicado: (2024)
Power-Law Decay Loss for Large Language Model Finetuning: A Theory Perspective
por: Shao, Jintian
Publicado: (2025)
por: Shao, Jintian
Publicado: (2025)
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and $\mathcal{O}(T)$ Complexity
por: Zou, Shihao, et al.
Publicado: (2025)
por: Zou, Shihao, et al.
Publicado: (2025)
From Complex Dynamics to DynFormer: Rethinking Transformers for PDEs
por: Lai, Pengyu, et al.
Publicado: (2026)
por: Lai, Pengyu, et al.
Publicado: (2026)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
por: Shan, Jiquan, et al.
Publicado: (2025)
por: Shan, Jiquan, et al.
Publicado: (2025)
Inter2Former: Dynamic Hybrid Attention for Efficient High-Precision Interactive
por: Huang, You, et al.
Publicado: (2025)
por: Huang, You, et al.
Publicado: (2025)
LATTE: Low-Precision Approximate Attention with Head-wise Trainable Threshold for Efficient Transformer
por: Wang, Jiing-Ping, et al.
Publicado: (2024)
por: Wang, Jiing-Ping, et al.
Publicado: (2024)
Multi-party Agent Relation Sampling for Multi-party Ad Hoc Teamwork
por: Zhang, Beiwen, et al.
Publicado: (2025)
por: Zhang, Beiwen, et al.
Publicado: (2025)
ReCQR: Incorporating conversational query rewriting to improve Multimodal Image Retrieval
por: Hu, Yuan, et al.
Publicado: (2026)
por: Hu, Yuan, et al.
Publicado: (2026)
MSDCC ‐Net: A Fine‐Scale Remote Sensing Extraction Method for Bare Surface Land in Rare Earth Mining Areas Based on Multi‐Scale Attention Mechanisms
por: Yingming Cai, et al.
Publicado: (2026)
por: Yingming Cai, et al.
Publicado: (2026)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
por: Setyawan, Novendra, et al.
Publicado: (2024)
por: Setyawan, Novendra, et al.
Publicado: (2024)
PosFormer: Recognizing Complex Handwritten Mathematical Expression with Position Forest Transformer
por: Guan, Tongkun, et al.
Publicado: (2024)
por: Guan, Tongkun, et al.
Publicado: (2024)
Scalable Complexity Control Facilitates Reasoning Ability of LLMs
por: Hang, Liangkai, et al.
Publicado: (2025)
por: Hang, Liangkai, et al.
Publicado: (2025)
CHAI: Clustered Head Attention for Efficient LLM Inference
por: Agarwal, Saurabh, et al.
Publicado: (2024)
por: Agarwal, Saurabh, et al.
Publicado: (2024)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
por: Shao, Jintian, et al.
Publicado: (2025)
por: Shao, Jintian, et al.
Publicado: (2025)
MatFormer: Nested Transformer for Elastic Inference
por: Devvrit, et al.
Publicado: (2023)
por: Devvrit, et al.
Publicado: (2023)
Singular Vectors of Attention Heads Align with Features
por: Franco, Gabriel, et al.
Publicado: (2026)
por: Franco, Gabriel, et al.
Publicado: (2026)
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection
por: Hua, Kai, et al.
Publicado: (2025)
por: Hua, Kai, et al.
Publicado: (2025)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
por: Huang, Baihe, et al.
Publicado: (2025)
por: Huang, Baihe, et al.
Publicado: (2025)
Causal Inference with Complex Treatments: A Survey
por: Wang, Yingrong, et al.
Publicado: (2024)
por: Wang, Yingrong, et al.
Publicado: (2024)
Subsystem Complexity and Measurements in Holography
por: Jian, Shao-Kai, et al.
Publicado: (2023)
por: Jian, Shao-Kai, et al.
Publicado: (2023)
Complexity Equals (Almost) Anything
por: Myers, Robert C., et al.
Publicado: (2024)
por: Myers, Robert C., et al.
Publicado: (2024)
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs
por: Ji, Tao, et al.
Publicado: (2025)
por: Ji, Tao, et al.
Publicado: (2025)
Comet: A Communication-efficient and Performant Approximation for Private Transformer Inference
por: Xu, Xiangrui, et al.
Publicado: (2024)
por: Xu, Xiangrui, et al.
Publicado: (2024)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
por: Kumar, Phani, et al.
Publicado: (2026)
por: Kumar, Phani, et al.
Publicado: (2026)
RTA-Former: Reverse Transformer Attention for Polyp Segmentation
por: Li, Zhikai, et al.
Publicado: (2024)
por: Li, Zhikai, et al.
Publicado: (2024)
SGFormer: Single-Layer Graph Transformers with Approximation-Free Linear Complexity
por: Wu, Qitian, et al.
Publicado: (2024)
por: Wu, Qitian, et al.
Publicado: (2024)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
por: Zhou, Jingbo, et al.
Publicado: (2026)
por: Zhou, Jingbo, et al.
Publicado: (2026)
Castling-ViT: Compressing Self-Attention via Switching Towards Linear-Angular Attention at Vision Transformer Inference
por: You, Haoran, et al.
Publicado: (2022)
por: You, Haoran, et al.
Publicado: (2022)
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
por: Setyawan, Novendra, et al.
Publicado: (2025)
por: Setyawan, Novendra, et al.
Publicado: (2025)
Complexity=Anything: Singularity Probes
por: Jørstad, Eivind, et al.
Publicado: (2023)
por: Jørstad, Eivind, et al.
Publicado: (2023)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
PolaFormer: Polarity-aware Linear Attention for Vision Transformers
por: Meng, Weikang, et al.
Publicado: (2025)
por: Meng, Weikang, et al.
Publicado: (2025)
Ejemplares similares
-
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits
por: Shao, Jintian, et al.
Publicado: (2025) -
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
por: Shao, Jintian, et al.
Publicado: (2025) -
EulerFormer: Sequential User Behavior Modeling with Complex Vector Attention
por: Tian, Zhen, et al.
Publicado: (2024) -
CipherFormer: Efficient Transformer Private Inference with Low Round Complexity
por: Wang, Weize, et al.
Publicado: (2024) -
RecurFormer: Not All Transformer Heads Need Self-Attention
por: Yan, Ruiqing, et al.
Publicado: (2024)