How Powerful are Decoder-Only Transformer Neural Models?
Fuente:
arXiv
Salvato in:
| Autore principale: | Roberts, Jesse |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On The Adaptation of Unlimiformer for Decoder-Only Transformers
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
di: Ahrabian, Kian, et al.
Pubblicazione: (2024)
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025)
di: Roy, Amartya, et al.
Pubblicazione: (2025)
Machine Translation with Large Language Models: Decoder Only vs. Encoder-Decoder
di: M., Abhinav P., et al.
Pubblicazione: (2024)
di: M., Abhinav P., et al.
Pubblicazione: (2024)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
di: Ziashahabi, Amir, et al.
Pubblicazione: (2025)
di: Ziashahabi, Amir, et al.
Pubblicazione: (2025)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)
Loss Landscape Degeneracy and Stagewise Development in Transformers
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
di: Hoogland, Jesse, et al.
Pubblicazione: (2024)
WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Pretrained Speech Recognition Transformers
di: Pandey, Akshat, et al.
Pubblicazione: (2025)
di: Pandey, Akshat, et al.
Pubblicazione: (2025)
LayerNorm Induces Recency Bias in Transformer Decoders
di: Kim, Junu, et al.
Pubblicazione: (2025)
di: Kim, Junu, et al.
Pubblicazione: (2025)
Dissociating Decodability and Causal Use in Bracket-Sequence Transformers
di: Sharma, Aryan, et al.
Pubblicazione: (2026)
di: Sharma, Aryan, et al.
Pubblicazione: (2026)
CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling
di: Baral, Aditeya, et al.
Pubblicazione: (2025)
di: Baral, Aditeya, et al.
Pubblicazione: (2025)
Faster Transformer Decoding: N-gram Masked Self-Attention
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
di: Chelba, Ciprian, et al.
Pubblicazione: (2020)
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
di: Adhikari, Rabin
Pubblicazione: (2025)
di: Adhikari, Rabin
Pubblicazione: (2025)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
di: Wang, Andrew Z., et al.
Pubblicazione: (2025)
Decoding Uncertainty: The Impact of Decoding Strategies for Uncertainty Estimation in Large Language Models
di: Hashimoto, Wataru, et al.
Pubblicazione: (2025)
di: Hashimoto, Wataru, et al.
Pubblicazione: (2025)
Plain Transformers Can be Powerful Graph Learners
di: Ma, Liheng, et al.
Pubblicazione: (2025)
di: Ma, Liheng, et al.
Pubblicazione: (2025)
Decoding Speculative Decoding
di: Yan, Minghao, et al.
Pubblicazione: (2024)
di: Yan, Minghao, et al.
Pubblicazione: (2024)
On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms
di: Trauger, Jacob, et al.
Pubblicazione: (2025)
di: Trauger, Jacob, et al.
Pubblicazione: (2025)
The Counting Power of Transformers
di: Sälzer, Marco, et al.
Pubblicazione: (2025)
di: Sälzer, Marco, et al.
Pubblicazione: (2025)
Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEs
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2025)
di: García-de-Herreros, Paloma, et al.
Pubblicazione: (2025)
PolarQuant: Leveraging Polar Transformation for Efficient Key Cache Quantization and Decoding Acceleration
di: Wu, Songhao, et al.
Pubblicazione: (2025)
di: Wu, Songhao, et al.
Pubblicazione: (2025)
The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices
di: Arias, Esteban Garces, et al.
Pubblicazione: (2026)
di: Arias, Esteban Garces, et al.
Pubblicazione: (2026)
Transformers meet Neural Algorithmic Reasoners
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
di: Bounsi, Wilfried, et al.
Pubblicazione: (2024)
A2SF: Accumulative Attention Scoring with Forgetting Factor for Token Pruning in Transformer Decoder
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
di: Jo, Hyun-rae, et al.
Pubblicazione: (2024)
Stability-Weighted Decoding for Diffusion Language Models
di: Wu, Yue, et al.
Pubblicazione: (2026)
di: Wu, Yue, et al.
Pubblicazione: (2026)
Learning to Decode Collaboratively with Multiple Language Models
di: Shen, Shannon Zejiang, et al.
Pubblicazione: (2024)
di: Shen, Shannon Zejiang, et al.
Pubblicazione: (2024)
Barriers to Universal Reasoning With Transformers (And How to Overcome Them)
di: Kraus, Oliver, et al.
Pubblicazione: (2026)
di: Kraus, Oliver, et al.
Pubblicazione: (2026)
Calibrating Large Language Models Using Their Generations Only
di: Ulmer, Dennis, et al.
Pubblicazione: (2024)
di: Ulmer, Dennis, et al.
Pubblicazione: (2024)
Accelerating Transformer Inference for Translation via Parallel Decoding
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
di: Santilli, Andrea, et al.
Pubblicazione: (2023)
On the Power of Convolution Augmented Transformer
di: Li, Mingchen, et al.
Pubblicazione: (2024)
di: Li, Mingchen, et al.
Pubblicazione: (2024)
Latent Chain-of-Thought? Decoding the Depth-Recurrent Transformer
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
di: Lu, Wenquan, et al.
Pubblicazione: (2025)
Only relative ranks matter in weight-clustered large language models
di: Aizpurua, Borja, et al.
Pubblicazione: (2026)
di: Aizpurua, Borja, et al.
Pubblicazione: (2026)
Training LLMs over Neurally Compressed Text
di: Lester, Brian, et al.
Pubblicazione: (2024)
di: Lester, Brian, et al.
Pubblicazione: (2024)
Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation
di: Ren, Liliang, et al.
Pubblicazione: (2025)
di: Ren, Liliang, et al.
Pubblicazione: (2025)
FlashDecoding++: Faster Large Language Model Inference on GPUs
di: Hong, Ke, et al.
Pubblicazione: (2023)
di: Hong, Ke, et al.
Pubblicazione: (2023)
Recurrent Drafter for Fast Speculative Decoding in Large Language Models
di: Cheng, Yunfei, et al.
Pubblicazione: (2024)
di: Cheng, Yunfei, et al.
Pubblicazione: (2024)
Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases
di: Carbonari, Valentina, et al.
Pubblicazione: (2025)
di: Carbonari, Valentina, et al.
Pubblicazione: (2025)
Decoding Decoded: Understanding Hyperparameter Effects in Open-Ended Text Generation
di: Arias, Esteban Garces, et al.
Pubblicazione: (2024)
di: Arias, Esteban Garces, et al.
Pubblicazione: (2024)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
di: Liu, James, et al.
Pubblicazione: (2024)
di: Liu, James, et al.
Pubblicazione: (2024)
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
Self-Rewarding PPO: Aligning Large Language Models with Demonstrations Only
di: Zhang, Qingru, et al.
Pubblicazione: (2025)
di: Zhang, Qingru, et al.
Pubblicazione: (2025)
Documenti analoghi
-
On The Adaptation of Unlimiformer for Decoder-Only Transformers
di: Ahrabian, Kian, et al.
Pubblicazione: (2024) -
Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
di: Roy, Amartya, et al.
Pubblicazione: (2025) -
Machine Translation with Large Language Models: Decoder Only vs. Encoder-Decoder
di: M., Abhinav P., et al.
Pubblicazione: (2024) -
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
di: Ziashahabi, Amir, et al.
Pubblicazione: (2025) -
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
di: Wiegand, Götz-Henrik, et al.
Pubblicazione: (2026)