Enregistré dans:
| Auteurs principaux: | Conzelmann, Alexander, Catalan-Tatjer, Albert, Liu, Shiwei |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2605.06366 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Training Dynamics Impact Post-Training Quantization Robustness
par: Catalan-Tatjer, Albert, et autres
Publié: (2025)
par: Catalan-Tatjer, Albert, et autres
Publié: (2025)
Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
par: Conzelmann, Alexander, et autres
Publié: (2025)
par: Conzelmann, Alexander, et autres
Publié: (2025)
Decentralized Task Offloading and Load-Balancing for Mobile Edge Computing in Dense Networks
par: Yahya, Mariam, et autres
Publié: (2024)
par: Yahya, Mariam, et autres
Publié: (2024)
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
par: Lu, Haiquan, et autres
Publié: (2024)
par: Lu, Haiquan, et autres
Publié: (2024)
Quantifying Error Propagation and Model Collapse in Diffusion Models
par: Khelifa, Nail B., et autres
Publié: (2026)
par: Khelifa, Nail B., et autres
Publié: (2026)
ActTail: Global Activation Sparsity in Large Language Models
par: Hou, Wenwen, et autres
Publié: (2026)
par: Hou, Wenwen, et autres
Publié: (2026)
Mix-LN: Unleashing the Power of Deeper Layers by Combining Pre-LN and Post-LN
par: Li, Pengxiang, et autres
Publié: (2024)
par: Li, Pengxiang, et autres
Publié: (2024)
dgMARK: Decoding-Guided Watermarking for Diffusion Language Models
par: Hong, Pyo Min, et autres
Publié: (2026)
par: Hong, Pyo Min, et autres
Publié: (2026)
Path-Dependent Denoising: A Non-Conservative Field Perspective on Order Collapse in Diffusion Language Models
par: Kim, Jeonseong
Publié: (2026)
par: Kim, Jeonseong
Publié: (2026)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
par: Zhang, Yi, et autres
Publié: (2025)
par: Zhang, Yi, et autres
Publié: (2025)
Dominating vs. Dominated: Generative Collapse in Diffusion Models
par: Jeong, Hayeon, et autres
Publié: (2025)
par: Jeong, Hayeon, et autres
Publié: (2025)
Layer Collapse Can be Induced by Unstructured Pruning
par: Liao, Zhu, et autres
Publié: (2024)
par: Liao, Zhu, et autres
Publié: (2024)
LayerCollapse: Adaptive compression of neural networks
par: Shabgahi, Soheil Zibakhsh, et autres
Publié: (2023)
par: Shabgahi, Soheil Zibakhsh, et autres
Publié: (2023)
LaCoOT: Layer Collapse through Optimal Transport
par: Quétu, Victor, et autres
Publié: (2024)
par: Quétu, Victor, et autres
Publié: (2024)
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
par: Wang, Keyu, et autres
Publié: (2025)
par: Wang, Keyu, et autres
Publié: (2025)
The Curse of Depth in Large Language Models
par: Sun, Wenfang, et autres
Publié: (2025)
par: Sun, Wenfang, et autres
Publié: (2025)
Strong Model Collapse
par: Dohmatob, Elvis, et autres
Publié: (2024)
par: Dohmatob, Elvis, et autres
Publié: (2024)
Collapse-Free Prototype Readout Layer for Transformer Encoders
par: Cirrincione, Giansalvo, et autres
Publié: (2026)
par: Cirrincione, Giansalvo, et autres
Publié: (2026)
Provably Learning Diffusion Models under the Manifold Hypothesis: Collapse and Refine
par: Huang, Wei, et autres
Publié: (2026)
par: Huang, Wei, et autres
Publié: (2026)
Language Generation with Replay: A Learning-Theoretic View of Model Collapse
par: Racca, Giorgio, et autres
Publié: (2026)
par: Racca, Giorgio, et autres
Publié: (2026)
Q-GaLore: Quantized GaLore with INT4 Projection and Layer-Adaptive Low-Rank Gradients
par: Zhang, Zhenyu, et autres
Publié: (2024)
par: Zhang, Zhenyu, et autres
Publié: (2024)
Diffusion Language Models for Speech Recognition
par: Naveriani, Davyd, et autres
Publié: (2026)
par: Naveriani, Davyd, et autres
Publié: (2026)
Model Unmerging: Making Your Models Unmergeable for Secure Model Sharing
par: Wang, Zihao, et autres
Publié: (2025)
par: Wang, Zihao, et autres
Publié: (2025)
Monitoring Neural Training with Topology: A Footprint-Predictable Collapse Index
par: Kalinowski, Alexander
Publié: (2026)
par: Kalinowski, Alexander
Publié: (2026)
When One Modality Rules Them All: Backdoor Modality Collapse in Multimodal Diffusion Models
par: Wang, Qitong, et autres
Publié: (2026)
par: Wang, Qitong, et autres
Publié: (2026)
Variational Language Concepts for Interpreting Foundation Language Models
par: Wang, Hengyi, et autres
Publié: (2024)
par: Wang, Hengyi, et autres
Publié: (2024)
Mind the Gap: a Spectral Analysis of Rank Collapse and Signal Propagation in Attention Layers
par: Saada, Thiziri Nait, et autres
Publié: (2024)
par: Saada, Thiziri Nait, et autres
Publié: (2024)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
par: Sanyal, Sunny, et autres
Publié: (2024)
par: Sanyal, Sunny, et autres
Publié: (2024)
Multi-modal Synthetic Data Training and Model Collapse: Insights from VLMs and Diffusion Models
par: Hu, Zizhao, et autres
Publié: (2025)
par: Hu, Zizhao, et autres
Publié: (2025)
DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
par: Li, Quanhao, et autres
Publié: (2026)
par: Li, Quanhao, et autres
Publié: (2026)
Scaling Embedding Layers in Language Models
par: Yu, Da, et autres
Publié: (2025)
par: Yu, Da, et autres
Publié: (2025)
LOST: Low-rank and Sparse Pre-training for Large Language Models
par: Li, Jiaxi, et autres
Publié: (2025)
par: Li, Jiaxi, et autres
Publié: (2025)
A Probabilistic Perspective on Model Collapse
par: Xu, Shirong, et autres
Publié: (2025)
par: Xu, Shirong, et autres
Publié: (2025)
On the Robustness of Neural Collapse and the Neural Collapse of Robustness
par: Su, Jingtong, et autres
Publié: (2023)
par: Su, Jingtong, et autres
Publié: (2023)
Reward Collapse in Aligning Large Language Models
par: Song, Ziang, et autres
Publié: (2023)
par: Song, Ziang, et autres
Publié: (2023)
How to Unlock Time Series Editing? Diffusion-Driven Approach with Multi-Grained Control
par: Yu, Hao, et autres
Publié: (2025)
par: Yu, Hao, et autres
Publié: (2025)
Addressing Representation Collapse in Vector Quantized Models with One Linear Layer
par: Zhu, Yongxin, et autres
Publié: (2024)
par: Zhu, Yongxin, et autres
Publié: (2024)
Simple Drop-in LoRA Conditioning on Attention Layers Will Improve Your Diffusion Model
par: Choi, Joo Young, et autres
Publié: (2024)
par: Choi, Joo Young, et autres
Publié: (2024)
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
par: Harrington, Anne, et autres
Publié: (2025)
par: Harrington, Anne, et autres
Publié: (2025)
L$^3$: Large Lookup Layers
par: Tseng, Albert, et autres
Publié: (2026)
par: Tseng, Albert, et autres
Publié: (2026)
Documents similaires
-
Training Dynamics Impact Post-Training Quantization Robustness
par: Catalan-Tatjer, Albert, et autres
Publié: (2025) -
Reducing Storage of Pretrained Neural Networks by Rate-Constrained Quantization and Entropy Coding
par: Conzelmann, Alexander, et autres
Publié: (2025) -
Decentralized Task Offloading and Load-Balancing for Mobile Edge Computing in Dense Networks
par: Yahya, Mariam, et autres
Publié: (2024) -
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
par: Lu, Haiquan, et autres
Publié: (2024) -
Quantifying Error Propagation and Model Collapse in Diffusion Models
par: Khelifa, Nail B., et autres
Publié: (2026)