Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wiegand, Götz-Henrik, Raichle, Lorena, Städeli, Rico, Hrycej, Tomas, Bermeitinger, Bernhard, Handschuh, Siegfried |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025)
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025)
Make Deep Networks Shallow Again
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023)
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023)
Efficient Neural Network Training via Subset Pretraining
von: Spörer, Jan, et al.
Veröffentlicht: (2024)
von: Spörer, Jan, et al.
Veröffentlicht: (2024)
Reducing the Transformer Architecture to a Minimum
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2024)
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2024)
Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task
von: Caillaut, Gaëtan, et al.
Veröffentlicht: (2024)
von: Caillaut, Gaëtan, et al.
Veröffentlicht: (2024)
Analyzing FOMC Minutes: Accuracy and Constraints of Language Models
von: Kim, Wonseong, et al.
Veröffentlicht: (2023)
von: Kim, Wonseong, et al.
Veröffentlicht: (2023)
Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
von: Kopiczko, Dawid J., et al.
Veröffentlicht: (2024)
Scaling Laws for Code: A More Data-Hungry Regime
von: Luo, Xianzhen, et al.
Veröffentlicht: (2025)
von: Luo, Xianzhen, et al.
Veröffentlicht: (2025)
Is Semantic Chunking Worth the Computational Cost?
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
von: Qu, Renyi, et al.
Veröffentlicht: (2024)
Scaling Laws for Speculative Decoding
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
von: Yan, Siyuan, et al.
Veröffentlicht: (2025)
Gated Tree Cross-Attention for Checkpoint-Compatible Syntax Injection in Decoder-Only LLMs
von: Gao, Xinyu, et al.
Veröffentlicht: (2026)
von: Gao, Xinyu, et al.
Veröffentlicht: (2026)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
von: Liu, James, et al.
Veröffentlicht: (2024)
von: Liu, James, et al.
Veröffentlicht: (2024)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
von: Schwethelm, Kristian, et al.
Veröffentlicht: (2026)
Encoder-Decoder or Decoder-Only? Revisiting Encoder-Decoder Large Language Model
von: Zhang, Biao, et al.
Veröffentlicht: (2025)
von: Zhang, Biao, et al.
Veröffentlicht: (2025)
How Do Decoder-Only LLMs Perceive Users? Rethinking Attention Masking for User Representation Learning
von: Yuan, Jiahao, et al.
Veröffentlicht: (2026)
von: Yuan, Jiahao, et al.
Veröffentlicht: (2026)
FLUX: Data Worth Training On
von: Gowtham, et al.
Veröffentlicht: (2026)
von: Gowtham, et al.
Veröffentlicht: (2026)
You Only Cache Once: Decoder-Decoder Architectures for Language Models
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
von: Sun, Yutao, et al.
Veröffentlicht: (2024)
A typology of marked-S languages
von: Handschuh, Corinna
Veröffentlicht: (2015)
von: Handschuh, Corinna
Veröffentlicht: (2015)
On The Adaptation of Unlimiformer for Decoder-Only Transformers
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
von: Ahrabian, Kian, et al.
Veröffentlicht: (2024)
What is Your Data Worth to GPT? LLM-Scale Data Valuation with Influence Functions
von: Choe, Sang Keun, et al.
Veröffentlicht: (2024)
von: Choe, Sang Keun, et al.
Veröffentlicht: (2024)
Post-Hoc Understanding of Metaphor Processing in Decoder-Only Language Models via Conditional Scale Entropy
von: Chakrabarti, Lawhori, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Lawhori, et al.
Veröffentlicht: (2026)
Tiny Aya: Bridging Scale and Multilingual Depth
von: Salamanca, Alejandro R., et al.
Veröffentlicht: (2026)
von: Salamanca, Alejandro R., et al.
Veröffentlicht: (2026)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
von: Huang, Hongzhi, et al.
Veröffentlicht: (2025)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
von: Papi, Sara, et al.
Veröffentlicht: (2026)
von: Papi, Sara, et al.
Veröffentlicht: (2026)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Tiny Recursive Reasoning with Mamba-2 Attention Hybrid
von: Wang, Wenlong, et al.
Veröffentlicht: (2026)
von: Wang, Wenlong, et al.
Veröffentlicht: (2026)
Leveraging Generative AI for Enhancing Domain-Driven Software Design
von: Wiegand, Götz-Henrik, et al.
Veröffentlicht: (2026)
von: Wiegand, Götz-Henrik, et al.
Veröffentlicht: (2026)
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
von: Ni, Jiliang, et al.
Veröffentlicht: (2025)
How Powerful are Decoder-Only Transformer Neural Models?
von: Roberts, Jesse
Veröffentlicht: (2023)
von: Roberts, Jesse
Veröffentlicht: (2023)
RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
von: Goldstein, Daniel, et al.
Veröffentlicht: (2025)
Hey, That's My Data! Token-Only Dataset Inference in Large Language Models
von: Xiong, Chen, et al.
Veröffentlicht: (2025)
von: Xiong, Chen, et al.
Veröffentlicht: (2025)
More Data, Fewer Diacritics: Scaling Arabic TTS
von: Musleh, Ahmed, et al.
Veröffentlicht: (2026)
von: Musleh, Ahmed, et al.
Veröffentlicht: (2026)
DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models
von: Ye, Xi, et al.
Veröffentlicht: (2026)
von: Ye, Xi, et al.
Veröffentlicht: (2026)
Linear-time Minimum Bayes Risk Decoding with Reference Aggregation
von: Vamvas, Jannis, et al.
Veröffentlicht: (2024)
von: Vamvas, Jannis, et al.
Veröffentlicht: (2024)
LLaMA based Punctuation Restoration With Forward Pass Only Decoding
von: Pang, Yutong, et al.
Veröffentlicht: (2024)
von: Pang, Yutong, et al.
Veröffentlicht: (2024)
Gender Disambiguation in Machine Translation: Diagnostic Evaluation in Decoder-Only Architectures
von: Manna, Chiara, et al.
Veröffentlicht: (2026)
von: Manna, Chiara, et al.
Veröffentlicht: (2026)
Machine Translation with Large Language Models: Decoder Only vs. Encoder-Decoder
von: M., Abhinav P., et al.
Veröffentlicht: (2024)
von: M., Abhinav P., et al.
Veröffentlicht: (2024)
Speculative Decoding Scaling Laws (SDSL): Throughput Optimization Made Simple
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
von: Bozorgkhoo, Amirhossein, et al.
Veröffentlicht: (2026)
Loss-to-Loss Prediction: Scaling Laws for All Datasets
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
von: Brandfonbrener, David, et al.
Veröffentlicht: (2024)
Reject Only Critical Tokens: Pivot-Aware Speculative Decoding
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
von: Ziashahabi, Amir, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Convexity-dependent Two-Phase Training Algorithm for Deep Neural Networks
von: Hrycej, Tomas, et al.
Veröffentlicht: (2025) -
Make Deep Networks Shallow Again
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2023) -
Efficient Neural Network Training via Subset Pretraining
von: Spörer, Jan, et al.
Veröffentlicht: (2024) -
Reducing the Transformer Architecture to a Minimum
von: Bermeitinger, Bernhard, et al.
Veröffentlicht: (2024) -
Scaling Laws of Decoder-Only Models on the Multilingual Machine Translation Task
von: Caillaut, Gaëtan, et al.
Veröffentlicht: (2024)