Towards smaller, faster decoder-only transformers: Architectural variants and their implications
Fuente:
arXiv
Guardado en:
| Autores principales: | Suresh, Sathya Krishnan, P, Shunmugapriya |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Generic Approach to Visualization of Time Series Data
por: Suresh, Sathya Krishnan, et al.
Publicado: (2022)
por: Suresh, Sathya Krishnan, et al.
Publicado: (2022)
Dynamic layer selection in decoder-only transformers
por: Glavas, Theodore, et al.
Publicado: (2024)
por: Glavas, Theodore, et al.
Publicado: (2024)
DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
por: Suresh, Sathya Krishnan, et al.
Publicado: (2024)
por: Suresh, Sathya Krishnan, et al.
Publicado: (2024)
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
por: Suresh, Sathya Krishnan, et al.
Publicado: (2025)
por: Suresh, Sathya Krishnan, et al.
Publicado: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
por: Loza, Andrew J., et al.
Publicado: (2025)
por: Loza, Andrew J., et al.
Publicado: (2025)
A decoder-only foundation model for time-series forecasting
por: Das, Abhimanyu, et al.
Publicado: (2023)
por: Das, Abhimanyu, et al.
Publicado: (2023)
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
por: Medapati, Sourabh, et al.
Publicado: (2025)
por: Medapati, Sourabh, et al.
Publicado: (2025)
The Effect of Architecture During Continual Learning
por: Hahn, Allyson, et al.
Publicado: (2026)
por: Hahn, Allyson, et al.
Publicado: (2026)
Less Memory Means smaller GPUs: Backpropagation with Compressed Activations
por: Barley, Daniel, et al.
Publicado: (2024)
por: Barley, Daniel, et al.
Publicado: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
por: Maiti, Soumi, et al.
Publicado: (2023)
por: Maiti, Soumi, et al.
Publicado: (2023)
Joint control variate for faster black-box variational inference
por: Wang, Xi, et al.
Publicado: (2022)
por: Wang, Xi, et al.
Publicado: (2022)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
por: Bartosh, Grigory, et al.
Publicado: (2026)
por: Bartosh, Grigory, et al.
Publicado: (2026)
LeanVec: Searching vectors faster by making them fit
por: Tepper, Mariano, et al.
Publicado: (2023)
por: Tepper, Mariano, et al.
Publicado: (2023)
Towards Improved Imbalance Robustness in Continual Multi-Label Learning with Dual Output Spiking Architecture (DOSA)
por: Mishra, Sourav, et al.
Publicado: (2024)
por: Mishra, Sourav, et al.
Publicado: (2024)
Divine Benevolence is an $x^2$: GLUs scale asymptotically faster than MLPs
por: Queiruga, Alejandro Francisco
Publicado: (2026)
por: Queiruga, Alejandro Francisco
Publicado: (2026)
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
por: Bao, Yuntai, et al.
Publicado: (2026)
por: Bao, Yuntai, et al.
Publicado: (2026)
Temporal convolutional and fusional transformer model with Bi-LSTM encoder-decoder for multi-time-window remaining useful life prediction
por: Pour, Mohamadreza Akbari, et al.
Publicado: (2025)
por: Pour, Mohamadreza Akbari, et al.
Publicado: (2025)
TabICLv2: A better, faster, scalable, and open tabular foundation model
por: Qu, Jingang, et al.
Publicado: (2026)
por: Qu, Jingang, et al.
Publicado: (2026)
Towards auditory attention decoding with noise-tagging: A pilot study
por: Scheppink, H. A., et al.
Publicado: (2024)
por: Scheppink, H. A., et al.
Publicado: (2024)
Enhancing the Inductive Biases of Graph Neural ODE for Modeling Dynamical Systems
por: Bishnoi, Suresh, et al.
Publicado: (2022)
por: Bishnoi, Suresh, et al.
Publicado: (2022)
Sketch and shift: a robust decoder for compressive clustering
por: Belhadji, Ayoub, et al.
Publicado: (2023)
por: Belhadji, Ayoub, et al.
Publicado: (2023)
Tractable Sharpness-Aware Learning of Probabilistic Circuits
por: Suresh, Hrithik, et al.
Publicado: (2025)
por: Suresh, Hrithik, et al.
Publicado: (2025)
CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation
por: Solís-García, Javier, et al.
Publicado: (2025)
por: Solís-García, Javier, et al.
Publicado: (2025)
From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training
por: Berner, Julius, et al.
Publicado: (2025)
por: Berner, Julius, et al.
Publicado: (2025)
Learning graph topology from metapopulation epidemic encoder-decoder
por: Li, Xin, et al.
Publicado: (2026)
por: Li, Xin, et al.
Publicado: (2026)
Forward-only Diffusion Probabilistic Models
por: Luo, Ziwei, et al.
Publicado: (2025)
por: Luo, Ziwei, et al.
Publicado: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
por: Yao, Zengwei, et al.
Publicado: (2023)
por: Yao, Zengwei, et al.
Publicado: (2023)
Self-Aligned Reward: Towards Effective and Efficient Reasoners
por: Han, Peixuan, et al.
Publicado: (2025)
por: Han, Peixuan, et al.
Publicado: (2025)
Efficient Nudged Elastic Band Method using Neural Network Bayesian Algorithm Execution
por: Kakhandiki, Pranav, et al.
Publicado: (2025)
por: Kakhandiki, Pranav, et al.
Publicado: (2025)
Not only where, But when: Temporal Scheduling for RLVR
por: Zhang, Jinghao, et al.
Publicado: (2026)
por: Zhang, Jinghao, et al.
Publicado: (2026)
Model soups need only one ingredient
por: Abdollahpoorrostam, Alireza, et al.
Publicado: (2026)
por: Abdollahpoorrostam, Alireza, et al.
Publicado: (2026)
Learning a Factorized Orthogonal Latent Space using Encoder-only Architecture for Fault Detection; An Alarm management perspective
por: Eivaghi, Vahid MohammadZadeh, et al.
Publicado: (2024)
por: Eivaghi, Vahid MohammadZadeh, et al.
Publicado: (2024)
Augmenting Offline Reinforcement Learning with State-only Interactions
por: Li, Shangzhe, et al.
Publicado: (2024)
por: Li, Shangzhe, et al.
Publicado: (2024)
Universally Converging Representations of Matter Across Scientific Foundation Models
por: Edamadaka, Sathya, et al.
Publicado: (2025)
por: Edamadaka, Sathya, et al.
Publicado: (2025)
Galactification: painting galaxies onto dark matter only simulations using a transformer-based model
por: Pandey, Shivam, et al.
Publicado: (2025)
por: Pandey, Shivam, et al.
Publicado: (2025)
EHRMamba: Towards Generalizable and Scalable Foundation Models for Electronic Health Records
por: Fallahpour, Adibvafa, et al.
Publicado: (2024)
por: Fallahpour, Adibvafa, et al.
Publicado: (2024)
DDIM sampling for Generative AIBIM, a faster intelligent structural design framework
por: He, Zhili, et al.
Publicado: (2024)
por: He, Zhili, et al.
Publicado: (2024)
Nonlocal operator learning for fMRI encoding and decoding tasks
por: Kramer, Andreas, et al.
Publicado: (2026)
por: Kramer, Andreas, et al.
Publicado: (2026)
Sequential decoder training for improved latent space dynamics identification
por: Anderson, William, et al.
Publicado: (2025)
por: Anderson, William, et al.
Publicado: (2025)
USDs: A universal stabilizer decoder framework using symmetry
por: Ohnishi, Hoshitaro, et al.
Publicado: (2026)
por: Ohnishi, Hoshitaro, et al.
Publicado: (2026)
Ejemplares similares
-
Generic Approach to Visualization of Time Series Data
por: Suresh, Sathya Krishnan, et al.
Publicado: (2022) -
Dynamic layer selection in decoder-only transformers
por: Glavas, Theodore, et al.
Publicado: (2024) -
DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
por: Suresh, Sathya Krishnan, et al.
Publicado: (2024) -
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
por: Suresh, Sathya Krishnan, et al.
Publicado: (2025) -
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
por: Loza, Andrew J., et al.
Publicado: (2025)