Towards smaller, faster decoder-only transformers: Architectural variants and their implications
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Suresh, Sathya Krishnan, P, Shunmugapriya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generic Approach to Visualization of Time Series Data
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2022)
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2022)
Dynamic layer selection in decoder-only transformers
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
von: Glavas, Theodore, et al.
Veröffentlicht: (2024)
DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2024)
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2024)
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2025)
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)
A decoder-only foundation model for time-series forecasting
von: Das, Abhimanyu, et al.
Veröffentlicht: (2023)
von: Das, Abhimanyu, et al.
Veröffentlicht: (2023)
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
von: Medapati, Sourabh, et al.
Veröffentlicht: (2025)
von: Medapati, Sourabh, et al.
Veröffentlicht: (2025)
The Effect of Architecture During Continual Learning
von: Hahn, Allyson, et al.
Veröffentlicht: (2026)
von: Hahn, Allyson, et al.
Veröffentlicht: (2026)
Less Memory Means smaller GPUs: Backpropagation with Compressed Activations
von: Barley, Daniel, et al.
Veröffentlicht: (2024)
von: Barley, Daniel, et al.
Veröffentlicht: (2024)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
von: Maiti, Soumi, et al.
Veröffentlicht: (2023)
Joint control variate for faster black-box variational inference
von: Wang, Xi, et al.
Veröffentlicht: (2022)
von: Wang, Xi, et al.
Veröffentlicht: (2022)
Forward-Learned Discrete Diffusion: Learning how to noise to denoise faster
von: Bartosh, Grigory, et al.
Veröffentlicht: (2026)
von: Bartosh, Grigory, et al.
Veröffentlicht: (2026)
LeanVec: Searching vectors faster by making them fit
von: Tepper, Mariano, et al.
Veröffentlicht: (2023)
von: Tepper, Mariano, et al.
Veröffentlicht: (2023)
Towards Improved Imbalance Robustness in Continual Multi-Label Learning with Dual Output Spiking Architecture (DOSA)
von: Mishra, Sourav, et al.
Veröffentlicht: (2024)
von: Mishra, Sourav, et al.
Veröffentlicht: (2024)
Divine Benevolence is an $x^2$: GLUs scale asymptotically faster than MLPs
von: Queiruga, Alejandro Francisco
Veröffentlicht: (2026)
von: Queiruga, Alejandro Francisco
Veröffentlicht: (2026)
Towards Steering without Sacrifice: Principled Training of Steering Vectors for Prompt-only Interventions
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
von: Bao, Yuntai, et al.
Veröffentlicht: (2026)
Temporal convolutional and fusional transformer model with Bi-LSTM encoder-decoder for multi-time-window remaining useful life prediction
von: Pour, Mohamadreza Akbari, et al.
Veröffentlicht: (2025)
von: Pour, Mohamadreza Akbari, et al.
Veröffentlicht: (2025)
TabICLv2: A better, faster, scalable, and open tabular foundation model
von: Qu, Jingang, et al.
Veröffentlicht: (2026)
von: Qu, Jingang, et al.
Veröffentlicht: (2026)
Towards auditory attention decoding with noise-tagging: A pilot study
von: Scheppink, H. A., et al.
Veröffentlicht: (2024)
von: Scheppink, H. A., et al.
Veröffentlicht: (2024)
Enhancing the Inductive Biases of Graph Neural ODE for Modeling Dynamical Systems
von: Bishnoi, Suresh, et al.
Veröffentlicht: (2022)
von: Bishnoi, Suresh, et al.
Veröffentlicht: (2022)
Sketch and shift: a robust decoder for compressive clustering
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
von: Belhadji, Ayoub, et al.
Veröffentlicht: (2023)
Tractable Sharpness-Aware Learning of Probabilistic Circuits
von: Suresh, Hrithik, et al.
Veröffentlicht: (2025)
von: Suresh, Hrithik, et al.
Veröffentlicht: (2025)
CoSTI: Consistency Models for (a faster) Spatio-Temporal Imputation
von: Solís-García, Javier, et al.
Veröffentlicht: (2025)
von: Solís-García, Javier, et al.
Veröffentlicht: (2025)
From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training
von: Berner, Julius, et al.
Veröffentlicht: (2025)
von: Berner, Julius, et al.
Veröffentlicht: (2025)
Learning graph topology from metapopulation epidemic encoder-decoder
von: Li, Xin, et al.
Veröffentlicht: (2026)
von: Li, Xin, et al.
Veröffentlicht: (2026)
Forward-only Diffusion Probabilistic Models
von: Luo, Ziwei, et al.
Veröffentlicht: (2025)
von: Luo, Ziwei, et al.
Veröffentlicht: (2025)
Zipformer: A faster and better encoder for automatic speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
von: Yao, Zengwei, et al.
Veröffentlicht: (2023)
Self-Aligned Reward: Towards Effective and Efficient Reasoners
von: Han, Peixuan, et al.
Veröffentlicht: (2025)
von: Han, Peixuan, et al.
Veröffentlicht: (2025)
Efficient Nudged Elastic Band Method using Neural Network Bayesian Algorithm Execution
von: Kakhandiki, Pranav, et al.
Veröffentlicht: (2025)
von: Kakhandiki, Pranav, et al.
Veröffentlicht: (2025)
Not only where, But when: Temporal Scheduling for RLVR
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
von: Zhang, Jinghao, et al.
Veröffentlicht: (2026)
Model soups need only one ingredient
von: Abdollahpoorrostam, Alireza, et al.
Veröffentlicht: (2026)
von: Abdollahpoorrostam, Alireza, et al.
Veröffentlicht: (2026)
Learning a Factorized Orthogonal Latent Space using Encoder-only Architecture for Fault Detection; An Alarm management perspective
von: Eivaghi, Vahid MohammadZadeh, et al.
Veröffentlicht: (2024)
von: Eivaghi, Vahid MohammadZadeh, et al.
Veröffentlicht: (2024)
Augmenting Offline Reinforcement Learning with State-only Interactions
von: Li, Shangzhe, et al.
Veröffentlicht: (2024)
von: Li, Shangzhe, et al.
Veröffentlicht: (2024)
Universally Converging Representations of Matter Across Scientific Foundation Models
von: Edamadaka, Sathya, et al.
Veröffentlicht: (2025)
von: Edamadaka, Sathya, et al.
Veröffentlicht: (2025)
Galactification: painting galaxies onto dark matter only simulations using a transformer-based model
von: Pandey, Shivam, et al.
Veröffentlicht: (2025)
von: Pandey, Shivam, et al.
Veröffentlicht: (2025)
EHRMamba: Towards Generalizable and Scalable Foundation Models for Electronic Health Records
von: Fallahpour, Adibvafa, et al.
Veröffentlicht: (2024)
von: Fallahpour, Adibvafa, et al.
Veröffentlicht: (2024)
DDIM sampling for Generative AIBIM, a faster intelligent structural design framework
von: He, Zhili, et al.
Veröffentlicht: (2024)
von: He, Zhili, et al.
Veröffentlicht: (2024)
Nonlocal operator learning for fMRI encoding and decoding tasks
von: Kramer, Andreas, et al.
Veröffentlicht: (2026)
von: Kramer, Andreas, et al.
Veröffentlicht: (2026)
Sequential decoder training for improved latent space dynamics identification
von: Anderson, William, et al.
Veröffentlicht: (2025)
von: Anderson, William, et al.
Veröffentlicht: (2025)
USDs: A universal stabilizer decoder framework using symmetry
von: Ohnishi, Hoshitaro, et al.
Veröffentlicht: (2026)
von: Ohnishi, Hoshitaro, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Generic Approach to Visualization of Time Series Data
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2022) -
Dynamic layer selection in decoder-only transformers
von: Glavas, Theodore, et al.
Veröffentlicht: (2024) -
DiaSynth: Synthetic Dialogue Generation Framework for Low Resource Dialogue Applications
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2024) -
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
von: Suresh, Sathya Krishnan, et al.
Veröffentlicht: (2025) -
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
von: Loza, Andrew J., et al.
Veröffentlicht: (2025)