Parcae: Scaling Laws For Stable Looped Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Prairie, Hayden, Novack, Zachary, Berg-Kirkpatrick, Taylor, Fu, Daniel Y. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
by: Long, Phillip, et al.
Published: (2024)
by: Long, Phillip, et al.
Published: (2024)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
by: Lee, Ivan, et al.
Published: (2025)
by: Lee, Ivan, et al.
Published: (2025)
Presto! Distilling Steps and Layers for Accelerating Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Low-Resource Guidance for Controllable Latent Audio Diffusion
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability
by: Lee, Ivan, et al.
Published: (2023)
by: Lee, Ivan, et al.
Published: (2023)
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
by: Long, Phillip, et al.
Published: (2026)
by: Long, Phillip, et al.
Published: (2026)
Upweighting Easy Samples in Fine-Tuning Mitigates Forgetting
by: Sanyal, Sunny, et al.
Published: (2025)
by: Sanyal, Sunny, et al.
Published: (2025)
Smaller Language Models are Better Black-box Machine-Generated Text Detectors
by: Mireshghallah, Niloofar, et al.
Published: (2023)
by: Mireshghallah, Niloofar, et al.
Published: (2023)
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
by: Schwethelm, Kristian, et al.
Published: (2026)
by: Schwethelm, Kristian, et al.
Published: (2026)
Continuous Diffusion Models Can Obey Formal Syntax
by: Kim, Jinwoo, et al.
Published: (2026)
by: Kim, Jinwoo, et al.
Published: (2026)
WildFX: A DAW-Powered Pipeline for In-the-Wild Audio FX Graph Modeling
by: Yang, Qihui, et al.
Published: (2025)
by: Yang, Qihui, et al.
Published: (2025)
Search Your Block Floating Point Scales!
by: Gupta, Tanmaey, et al.
Published: (2026)
by: Gupta, Tanmaey, et al.
Published: (2026)
Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
by: Novack, Zachary, et al.
Published: (2026)
by: Novack, Zachary, et al.
Published: (2026)
Studying the Soupability of Documents in State Space Models
by: Jafari, Yasaman, et al.
Published: (2025)
by: Jafari, Yasaman, et al.
Published: (2025)
Scaling Laws for Differentially Private Language Models
by: McKenna, Ryan, et al.
Published: (2025)
by: McKenna, Ryan, et al.
Published: (2025)
Generalizing Scaling Laws for Dense and Sparse Large Language Models
by: Hossain, Md Arafat, et al.
Published: (2025)
by: Hossain, Md Arafat, et al.
Published: (2025)
Fast Text-to-Audio Generation with Adversarial Post-Training
by: Novack, Zachary, et al.
Published: (2025)
by: Novack, Zachary, et al.
Published: (2025)
Constrained Sampling for Language Models Should Be Easy: An MCMC Perspective
by: Gonzalez, Emmanuel Anaya, et al.
Published: (2025)
by: Gonzalez, Emmanuel Anaya, et al.
Published: (2025)
Neural Scaling Laws of Deep ReLU and Deep Operator Network: A Theoretical Study
by: Liu, Hao, et al.
Published: (2024)
by: Liu, Hao, et al.
Published: (2024)
Optical Context Compression Is Just (Bad) Autoencoding
by: Lee, Ivan Yee, et al.
Published: (2025)
by: Lee, Ivan Yee, et al.
Published: (2025)
BACHI: Boundary-Aware Symbolic Chord Recognition Through Masked Iterative Decoding on Pop and Classical Music
by: Yao, Mingyang, et al.
Published: (2025)
by: Yao, Mingyang, et al.
Published: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
Constrained Adaptive Rejection Sampling
by: Parys, Paweł, et al.
Published: (2025)
by: Parys, Paweł, et al.
Published: (2025)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Scaling Laws for Multilingual Language Models
by: He, Yifei, et al.
Published: (2024)
by: He, Yifei, et al.
Published: (2024)
Alt-Text with Context: Improving Accessibility for Images on Twitter
by: Srivatsan, Nikita, et al.
Published: (2023)
by: Srivatsan, Nikita, et al.
Published: (2023)
Pretraining Scaling Laws for Generative Evaluations of Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
by: Chen, Danlu, et al.
Published: (2026)
by: Chen, Danlu, et al.
Published: (2026)
Sparse Layers are Critical to Scaling Looped Language Models
by: Lee, Ryan, et al.
Published: (2026)
by: Lee, Ryan, et al.
Published: (2026)
Scaling Laws for Precision
by: Kumar, Tanishq, et al.
Published: (2024)
by: Kumar, Tanishq, et al.
Published: (2024)
Learning the Error Patterns of Language Models
by: Kim, Jinwoo, et al.
Published: (2026)
by: Kim, Jinwoo, et al.
Published: (2026)
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
by: Subramanyam, Anirudh, et al.
Published: (2025)
by: Subramanyam, Anirudh, et al.
Published: (2025)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
by: Goyal, Sachin, et al.
Published: (2024)
by: Goyal, Sachin, et al.
Published: (2024)
Scaling Laws for Upcycling Mixture-of-Experts Language Models
by: Liew, Seng Pei, et al.
Published: (2025)
by: Liew, Seng Pei, et al.
Published: (2025)
Scaling Laws for Discriminative Classification in Large Language Models
by: Wyatte, Dean, et al.
Published: (2024)
by: Wyatte, Dean, et al.
Published: (2024)
Grammar-Aligned Decoding
by: Park, Kanghee, et al.
Published: (2024)
by: Park, Kanghee, et al.
Published: (2024)
Improving Generalization of Speech Separation in Real-World Scenarios: Strategies in Simulation, Optimization, and Evaluation
by: Chen, Ke, et al.
Published: (2024)
by: Chen, Ke, et al.
Published: (2024)
Similar Items
-
PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing
by: Long, Phillip, et al.
Published: (2024) -
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024) -
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025) -
DITTO: Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024) -
Readability $\ne$ Learnability: Rethinking the Role of Simplicity in Training Small Language Models
by: Lee, Ivan, et al.
Published: (2025)