Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Sernau, Luke |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Time Matters: Scaling Laws for Any Budget
von: Inbar, Itay, et al.
Veröffentlicht: (2024)
von: Inbar, Itay, et al.
Veröffentlicht: (2024)
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
von: Cooper, A. Feder, et al.
Veröffentlicht: (2024)
von: Cooper, A. Feder, et al.
Veröffentlicht: (2024)
MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
von: Krohn-Grimberghe, Artus
Veröffentlicht: (2026)
von: Krohn-Grimberghe, Artus
Veröffentlicht: (2026)
All Random Features Representations are Equivalent
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
von: Sernau, Luke, et al.
Veröffentlicht: (2024)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
Data-Aware Random Feature Kernel for Transformers
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
von: Farzam, Amirhossein, et al.
Veröffentlicht: (2026)
One LR Doesn't Fit All: Heavy-Tail Guided Layerwise Learning Rates for LLMs
von: He, Di, et al.
Veröffentlicht: (2026)
von: He, Di, et al.
Veröffentlicht: (2026)
Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences
von: Falahati, Ali, et al.
Veröffentlicht: (2026)
von: Falahati, Ali, et al.
Veröffentlicht: (2026)
Global Convergence and Rich Feature Learning in $L$-Layer Infinite-Width Neural Networks under $μ$P Parametrization
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
von: Chen, Zixiang, et al.
Veröffentlicht: (2025)
How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhai, Zhiyuan, et al.
Veröffentlicht: (2026)
Semantics at an Angle: When Cosine Similarity Works Until It Doesn't
von: You, Kisung
Veröffentlicht: (2025)
von: You, Kisung
Veröffentlicht: (2025)
Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2025)
von: Dang, Quy-Anh, et al.
Veröffentlicht: (2025)
Think Before You Act: Decision Transformers with Working Memory
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
von: Kang, Jikun, et al.
Veröffentlicht: (2023)
The Graphon Limit Hypothesis: Understanding Neural Network Pruning via Infinite Width Analysis
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis
von: Gong, Shuzhi, et al.
Veröffentlicht: (2026)
von: Gong, Shuzhi, et al.
Veröffentlicht: (2026)
Les Houches Lectures on Deep Learning at Large & Infinite Width
von: Bahri, Yasaman, et al.
Veröffentlicht: (2023)
von: Bahri, Yasaman, et al.
Veröffentlicht: (2023)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Feature Learning Dynamics in Infinite-Depth Neural Networks
von: Yao, Zihan, et al.
Veröffentlicht: (2025)
von: Yao, Zihan, et al.
Veröffentlicht: (2025)
On the Diminishing Returns of Width for Continual Learning
von: Guha, Etash, et al.
Veröffentlicht: (2024)
von: Guha, Etash, et al.
Veröffentlicht: (2024)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
von: He, Qianxi, et al.
Veröffentlicht: (2025)
von: He, Qianxi, et al.
Veröffentlicht: (2025)
ResNets Are Deeper Than You Think
von: Mehmeti-Göpel, Christian H. X. Ali, et al.
Veröffentlicht: (2025)
von: Mehmeti-Göpel, Christian H. X. Ali, et al.
Veröffentlicht: (2025)
What LLMs Think When You Don't Tell Them What to Think About?
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
von: Kwon, Yongchan, et al.
Veröffentlicht: (2026)
Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
von: Lawrence, Nathan P., et al.
Veröffentlicht: (2025)
von: Lawrence, Nathan P., et al.
Veröffentlicht: (2025)
Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
von: Zhao, Yike, et al.
Veröffentlicht: (2026)
von: Zhao, Yike, et al.
Veröffentlicht: (2026)
Why Does RLAIF Work At All?
von: Young, Robin
Veröffentlicht: (2026)
von: Young, Robin
Veröffentlicht: (2026)
The Infinite-Dimensional Nature of Spectroscopy and Why Models Succeed, Fail, and Mislead
von: Michelucci, Umberto, et al.
Veröffentlicht: (2026)
von: Michelucci, Umberto, et al.
Veröffentlicht: (2026)
Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
von: Juzek, Tom S., et al.
Veröffentlicht: (2024)
A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)
von: Nayak, Nihal V., et al.
Veröffentlicht: (2026)
von: Nayak, Nihal V., et al.
Veröffentlicht: (2026)
On the Infinite Width and Depth Limits of Predictive Coding Networks
von: Innocenti, Francesco, et al.
Veröffentlicht: (2026)
von: Innocenti, Francesco, et al.
Veröffentlicht: (2026)
The Pragmatic Frames of Spurious Correlations in Machine Learning: Interpreting How and Why They Matter
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
von: Bell, Samuel J., et al.
Veröffentlicht: (2024)
Virtual Width Networks
von: Seed, et al.
Veröffentlicht: (2025)
von: Seed, et al.
Veröffentlicht: (2025)
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
von: Li, Weiqi, et al.
Veröffentlicht: (2025)
von: Li, Weiqi, et al.
Veröffentlicht: (2025)
Position: Model Collapse Does Not Mean What You Think
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
von: Schaeffer, Rylan, et al.
Veröffentlicht: (2025)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
von: Clark, Tyler, et al.
Veröffentlicht: (2025)
von: Clark, Tyler, et al.
Veröffentlicht: (2025)
When More Data Doesn't Help: Limits of Adaptation in Multitask Learning
von: Hanneke, Steve, et al.
Veröffentlicht: (2026)
von: Hanneke, Steve, et al.
Veröffentlicht: (2026)
Bayesian Elicitation with LLMs: Model Size Helps, Extra "Reasoning" Doesn't Always
von: Hobor, Luka, et al.
Veröffentlicht: (2026)
von: Hobor, Luka, et al.
Veröffentlicht: (2026)
Adaptive Width Neural Networks
von: Errica, Federico, et al.
Veröffentlicht: (2025)
von: Errica, Federico, et al.
Veröffentlicht: (2025)
Why Representation Engineering Works: A Theoretical and Empirical Study in Vision-Language Models
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
von: Tian, Bowei, et al.
Veröffentlicht: (2025)
ThinkTrap: Denial-of-Service Attacks against Black-box LLM Services via Infinite Thinking
von: Li, Yunzhe, et al.
Veröffentlicht: (2025)
von: Li, Yunzhe, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Time Matters: Scaling Laws for Any Budget
von: Inbar, Itay, et al.
Veröffentlicht: (2024) -
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
von: Cooper, A. Feder, et al.
Veröffentlicht: (2024) -
MedCalc-Bench Doesn't Measure What You Think: A Benchmark Audit and the Case for Open-Book Evaluation
von: Krohn-Grimberghe, Artus
Veröffentlicht: (2026) -
All Random Features Representations are Equivalent
von: Sernau, Luke, et al.
Veröffentlicht: (2024) -
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
von: Yu, Zony, et al.
Veröffentlicht: (2025)