How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Schwethelm, Kristian, Rueckert, Daniel, Kaissis, Georgios |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
by: Kaiser, Johannes, et al.
Published: (2025)
by: Kaiser, Johannes, et al.
Published: (2025)
ChEX: Interactive Localization and Region Description in Chest X-rays
by: Müller, Philip, et al.
Published: (2024)
by: Müller, Philip, et al.
Published: (2024)
Differentially Private Active Learning: Balancing Effective Data Selection and Privacy
by: Schwethelm, Kristian, et al.
Published: (2024)
by: Schwethelm, Kristian, et al.
Published: (2024)
On Arbitrary Predictions from Equally Valid Models
by: Lockfisch, Sarah, et al.
Published: (2025)
by: Lockfisch, Sarah, et al.
Published: (2025)
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
by: Szep, Marton, et al.
Published: (2026)
by: Szep, Marton, et al.
Published: (2026)
Weakly Supervised Object Detection in Chest X-Rays with Differentiable ROI Proposal Networks and Soft ROI Pooling
by: Müller, Philip, et al.
Published: (2024)
by: Müller, Philip, et al.
Published: (2024)
Kernel Normalized Convolutional Networks
by: Nasirigerdeh, Reza, et al.
Published: (2022)
by: Nasirigerdeh, Reza, et al.
Published: (2022)
Is More Data Worth the Cost? Dataset Scaling Laws in a Tiny Attention-Only Decoder
by: Wiegand, Götz-Henrik, et al.
Published: (2026)
by: Wiegand, Götz-Henrik, et al.
Published: (2026)
Do Large Language Models Know How Much They Know?
by: Prato, Gabriele, et al.
Published: (2025)
by: Prato, Gabriele, et al.
Published: (2025)
Gradient-Weight Alignment as a Train-Time Proxy for Generalization in Classification Tasks
by: Hölzl, Florian A., et al.
Published: (2025)
by: Hölzl, Florian A., et al.
Published: (2025)
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
by: Kohli, Harsh, et al.
Published: (2026)
by: Kohli, Harsh, et al.
Published: (2026)
Incentivising the federation: gradient-based metrics for data selection and valuation in private decentralised training
by: Usynin, Dmitrii, et al.
Published: (2023)
by: Usynin, Dmitrii, et al.
Published: (2023)
Efficient Parallel Samplers for Recurrent-Depth Models and Their Connection to Diffusion Language Models
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
by: Bilmes, Jeff A., et al.
Published: (2026)
by: Bilmes, Jeff A., et al.
Published: (2026)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
by: Knupp, Jonas, et al.
Published: (2026)
by: Knupp, Jonas, et al.
Published: (2026)
Sparse Layers are Critical to Scaling Looped Language Models
by: Lee, Ryan, et al.
Published: (2026)
by: Lee, Ryan, et al.
Published: (2026)
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling
by: Huang, Hongzhi, et al.
Published: (2025)
by: Huang, Hongzhi, et al.
Published: (2025)
BertaQA: How Much Do Language Models Know About Local Culture?
by: Etxaniz, Julen, et al.
Published: (2024)
by: Etxaniz, Julen, et al.
Published: (2024)
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
by: Geiping, Jonas, et al.
Published: (2025)
by: Geiping, Jonas, et al.
Published: (2025)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Scaling Laws for Discriminative Classification in Large Language Models
by: Wyatte, Dean, et al.
Published: (2024)
by: Wyatte, Dean, et al.
Published: (2024)
Visual Privacy Auditing with Diffusion Models
by: Schwethelm, Kristian, et al.
Published: (2024)
by: Schwethelm, Kristian, et al.
Published: (2024)
BitDelta: Your Fine-Tune May Only Be Worth One Bit
by: Liu, James, et al.
Published: (2024)
by: Liu, James, et al.
Published: (2024)
Machine Unlearning for Medical Imaging
by: Nasirigerdeh, Reza, et al.
Published: (2024)
by: Nasirigerdeh, Reza, et al.
Published: (2024)
Scaling Laws for Post Training Quantized Large Language Models
by: Xu, Zifei, et al.
Published: (2024)
by: Xu, Zifei, et al.
Published: (2024)
Scaling Law for Language Models Training Considering Batch Size
by: Shuai, Xian, et al.
Published: (2024)
by: Shuai, Xian, et al.
Published: (2024)
Scaling Laws for Downstream Task Performance of Large Language Models
by: Isik, Berivan, et al.
Published: (2024)
by: Isik, Berivan, et al.
Published: (2024)
Fully Hyperbolic Convolutional Neural Networks for Computer Vision
by: Bdeir, Ahmad, et al.
Published: (2023)
by: Bdeir, Ahmad, et al.
Published: (2023)
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
by: Sardana, Nikhil, et al.
Published: (2023)
by: Sardana, Nikhil, et al.
Published: (2023)
Exploring Scaling Laws for Local SGD in Large Language Model Training
by: He, Qiaozhi, et al.
Published: (2024)
by: He, Qiaozhi, et al.
Published: (2024)
Can Language Models Discover Scaling Laws?
by: Lin, Haowei, et al.
Published: (2025)
by: Lin, Haowei, et al.
Published: (2025)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
by: Rathore, Darshita, et al.
Published: (2025)
by: Rathore, Darshita, et al.
Published: (2025)
Problems with Chinchilla Approach 2: Systematic Biases in IsoFLOP Parabola Fits
by: Czech, Eric, et al.
Published: (2026)
by: Czech, Eric, et al.
Published: (2026)
Counterfactual Influence as a Distributional Quantity
by: Meeus, Matthieu, et al.
Published: (2025)
by: Meeus, Matthieu, et al.
Published: (2025)
How to Upscale Neural Networks with Scaling Law? A Survey and Practical Guidelines
by: Sengupta, Ayan, et al.
Published: (2025)
by: Sengupta, Ayan, et al.
Published: (2025)
How Much Can We Forget about Data Contamination?
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
Relative-Based Scaling Law for Neural Language Models
by: Yue, Baoqing, et al.
Published: (2025)
by: Yue, Baoqing, et al.
Published: (2025)
Observational Scaling Laws and the Predictability of Language Model Performance
by: Ruan, Yangjun, et al.
Published: (2024)
by: Ruan, Yangjun, et al.
Published: (2024)
Similar Items
-
Laplace Sample Information: Data Informativeness Through a Bayesian Lens
by: Kaiser, Johannes, et al.
Published: (2025) -
ChEX: Interactive Localization and Region Description in Chest X-rays
by: Müller, Philip, et al.
Published: (2024) -
Differentially Private Active Learning: Balancing Effective Data Selection and Privacy
by: Schwethelm, Kristian, et al.
Published: (2024) -
On Arbitrary Predictions from Equally Valid Models
by: Lockfisch, Sarah, et al.
Published: (2025) -
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
by: Szep, Marton, et al.
Published: (2026)