On the Power of Convolution Augmented Transformer
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Mingchen, Zhang, Xuechen, Huang, Yixiao, Oymak, Samet |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BrainTransformers: SNN-LLM
par: Tang, Zhengzheng, et autres
Publié: (2024)
par: Tang, Zhengzheng, et autres
Publié: (2024)
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
par: Csordás, Róbert, et autres
Publié: (2023)
par: Csordás, Róbert, et autres
Publié: (2023)
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
par: Wang, Hanrui, et autres
Publié: (2020)
par: Wang, Hanrui, et autres
Publié: (2020)
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
par: Reda, Eslam, et autres
Publié: (2026)
par: Reda, Eslam, et autres
Publié: (2026)
Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language Model
par: Tang, Kaiwen, et autres
Publié: (2024)
par: Tang, Kaiwen, et autres
Publié: (2024)
SpikeLLM: Scaling up Spiking Neural Network to Large Language Models via Saliency-based Spiking
par: Xing, Xingrun, et autres
Publié: (2024)
par: Xing, Xingrun, et autres
Publié: (2024)
SpikingSSMs: Learning Long Sequences with Sparse and Parallel Spiking State Space Models
par: Shen, Shuaijie, et autres
Publié: (2024)
par: Shen, Shuaijie, et autres
Publié: (2024)
SpikeLM: Towards General Spike-Driven Language Modeling via Elastic Bi-Spiking Mechanisms
par: Xing, Xingrun, et autres
Publié: (2024)
par: Xing, Xingrun, et autres
Publié: (2024)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
par: Majumdar, Somshubra, et autres
Publié: (2024)
par: Majumdar, Somshubra, et autres
Publié: (2024)
SpikeGPT: Generative Pre-trained Language Model with Spiking Neural Networks
par: Zhu, Rui-Jie, et autres
Publié: (2023)
par: Zhu, Rui-Jie, et autres
Publié: (2023)
Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for Large Language Models
par: Dong, Peijie, et autres
Publié: (2024)
par: Dong, Peijie, et autres
Publié: (2024)
GridPE: Unifying Positional Encoding in Transformers with a Grid Cell-Inspired Framework
par: Li, Boyang, et autres
Publié: (2024)
par: Li, Boyang, et autres
Publié: (2024)
EvoX: Meta-Evolution for Automated Discovery
par: Liu, Shu, et autres
Publié: (2026)
par: Liu, Shu, et autres
Publié: (2026)
GLU Attention Improve Transformer
par: Wang, Zehao
Publié: (2025)
par: Wang, Zehao
Publié: (2025)
Hysteresis Activation Function for Efficient Inference
par: Kimhi, Moshe, et autres
Publié: (2024)
par: Kimhi, Moshe, et autres
Publié: (2024)
Large Language Models for Tuning Evolution Strategies
par: Kramer, Oliver
Publié: (2024)
par: Kramer, Oliver
Publié: (2024)
An enhanced Teaching-Learning-Based Optimization (TLBO) with Grey Wolf Optimizer (GWO) for text feature selection and clustering
par: Azarshab, Mahsa, et autres
Publié: (2024)
par: Azarshab, Mahsa, et autres
Publié: (2024)
Improving Sequence-to-Sequence Models for Abstractive Text Summarization Using Meta Heuristic Approaches
par: Saxena, Aditya, et autres
Publié: (2024)
par: Saxena, Aditya, et autres
Publié: (2024)
B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory
par: Zancato, Luca, et autres
Publié: (2024)
par: Zancato, Luca, et autres
Publié: (2024)
Neural Information Organizing and Processing -- Neural Machines
par: Petrila, Iosif Iulian
Publié: (2024)
par: Petrila, Iosif Iulian
Publié: (2024)
EvolKV: Evolutionary KV Cache Compression for LLM Inference
par: Yu, Bohan, et autres
Publié: (2025)
par: Yu, Bohan, et autres
Publié: (2025)
Decomposing Evolutionary Mixture-of-LoRA Architectures: The Routing Lever, the Lifecycle Penalty, and a Substrate-Conditional Boundary
par: Kumaresan, Ramchand
Publié: (2026)
par: Kumaresan, Ramchand
Publié: (2026)
An In-depth Walkthrough on Evolution of Neural Machine Translation
par: Jagtap, Rohan, et autres
Publié: (2020)
par: Jagtap, Rohan, et autres
Publié: (2020)
Large Language Models Suffer From Their Own Output: An Analysis of the Self-Consuming Training Loop
par: Briesch, Martin, et autres
Publié: (2023)
par: Briesch, Martin, et autres
Publié: (2023)
Pre-trained Language Models Learn Remarkably Accurate Representations of Numbers
par: Kadlčík, Marek, et autres
Publié: (2025)
par: Kadlčík, Marek, et autres
Publié: (2025)
AP-BMM: Approximating Capability-Cost Pareto Sets of LLMs via Asynchronous Prior-Guided Bayesian Model Merging
par: Chen, Kesheng, et autres
Publié: (2025)
par: Chen, Kesheng, et autres
Publié: (2025)
Intelligent Neural Networks: From Layered Architectures to Graph-Organized Intelligence
par: Salomon, Antoine
Publié: (2025)
par: Salomon, Antoine
Publié: (2025)
ComplicaCode: Enhancing Disease Complication Detection in Electronic Health Records through ICD Path Generation
par: Zhou, Xiaofan
Publié: (2023)
par: Zhou, Xiaofan
Publié: (2023)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
par: Gong, Shuzhi, et autres
Publié: (2026)
par: Gong, Shuzhi, et autres
Publié: (2026)
Improving Language Plasticity via Pretraining with Active Forgetting
par: Chen, Yihong, et autres
Publié: (2023)
par: Chen, Yihong, et autres
Publié: (2023)
Minimal Convolutional RNNs Accelerate Spatiotemporal Learning
par: Horuz, Coşku Can, et autres
Publié: (2025)
par: Horuz, Coşku Can, et autres
Publié: (2025)
A Transformer-based Neural Architecture Search Method
par: Wang, Shang, et autres
Publié: (2025)
par: Wang, Shang, et autres
Publié: (2025)
NOBLE: Accelerating Transformers with Nonlinear Low-Rank Branches
par: Smith, Ethan
Publié: (2026)
par: Smith, Ethan
Publié: (2026)
Generalized Dynamic Brain Functional Connectivity Based on Random Convolutions
par: Duan, Yongjie, et autres
Publié: (2024)
par: Duan, Yongjie, et autres
Publié: (2024)
Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
par: Munkhdalai, Tsendsuren, et autres
Publié: (2024)
par: Munkhdalai, Tsendsuren, et autres
Publié: (2024)
Position as Probability: Self-Supervised Transformers that Think Past Their Training for Length Extrapolation
par: Lee, Philip Heejun
Publié: (2025)
par: Lee, Philip Heejun
Publié: (2025)
A Comprehensive Survey of Convolutions in Deep Learning: Applications, Challenges, and Future Trends
par: Younesi, Abolfazl, et autres
Publié: (2024)
par: Younesi, Abolfazl, et autres
Publié: (2024)
Predicting the Stay Length of Patients in Hospitals using Convolutional Gated Recurrent Deep Learning Model
par: Neshat, Mehdi, et autres
Publié: (2024)
par: Neshat, Mehdi, et autres
Publié: (2024)
GT-SNT: A Linear-Time Transformer for Large-Scale Graphs via Spiking Node Tokenization
par: Zhang, Huizhe, et autres
Publié: (2025)
par: Zhang, Huizhe, et autres
Publié: (2025)
Interpretable Solutions for Breast Cancer Diagnosis with Grammatical Evolution and Data Augmentation
par: Hasan, Yumnah, et autres
Publié: (2024)
par: Hasan, Yumnah, et autres
Publié: (2024)
Documents similaires
-
BrainTransformers: SNN-LLM
par: Tang, Zhengzheng, et autres
Publié: (2024) -
SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
par: Csordás, Róbert, et autres
Publié: (2023) -
HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
par: Wang, Hanrui, et autres
Publié: (2020) -
A Hormone-inspired Emotion Layer for Transformer language models (HELT)
par: Reda, Eslam, et autres
Publié: (2026) -
Sorbet: A Neuromorphic Hardware-Compatible Transformer-Based Spiking Language Model
par: Tang, Kaiwen, et autres
Publié: (2024)