How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
Fuente:
arXiv
Guardado en:
| Autores principales: | Littwin, Etai, Saremi, Omid, Advani, Madhu, Thilak, Vimal, Nakkiran, Preetum, Huang, Chen, Susskind, Joshua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vanishing Gradients in Reinforcement Finetuning of Language Models
por: Razin, Noam, et al.
Publicado: (2023)
por: Razin, Noam, et al.
Publicado: (2023)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
por: Huang, Chen, et al.
Publicado: (2026)
por: Huang, Chen, et al.
Publicado: (2026)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
por: Li, Xianhang, et al.
Publicado: (2025)
por: Li, Xianhang, et al.
Publicado: (2025)
Annotations Mitigate Post-Training Mode Collapse
por: Springer, Jacob Mitchell, et al.
Publicado: (2026)
por: Springer, Jacob Mitchell, et al.
Publicado: (2026)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
por: Littwin, Etai, et al.
Publicado: (2024)
por: Littwin, Etai, et al.
Publicado: (2024)
Step-by-Step Diffusion: An Elementary Tutorial
por: Nakkiran, Preetum, et al.
Publicado: (2024)
por: Nakkiran, Preetum, et al.
Publicado: (2024)
When can transformers reason with abstract symbols?
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
por: Boix-Adsera, Enric, et al.
Publicado: (2023)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
por: Malach, Eran, et al.
Publicado: (2025)
por: Malach, Eran, et al.
Publicado: (2025)
Mechanisms of Projective Composition of Diffusion Models
por: Bradley, Arwen, et al.
Publicado: (2025)
por: Bradley, Arwen, et al.
Publicado: (2025)
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
por: Fu, Yicheng, et al.
Publicado: (2024)
por: Fu, Yicheng, et al.
Publicado: (2024)
Classifier-Free Guidance is a Predictor-Corrector
por: Bradley, Arwen, et al.
Publicado: (2024)
por: Bradley, Arwen, et al.
Publicado: (2024)
Distillation Scaling Laws
por: Busbridge, Dan, et al.
Publicado: (2025)
por: Busbridge, Dan, et al.
Publicado: (2025)
Composition and Control with Distilled Energy Diffusion Models and Sequential Monte Carlo
por: Thornton, James, et al.
Publicado: (2025)
por: Thornton, James, et al.
Publicado: (2025)
When is Multicalibration Post-Processing Necessary?
por: Hansen, Dutch, et al.
Publicado: (2024)
por: Hansen, Dutch, et al.
Publicado: (2024)
Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
por: Abnar, Samira, et al.
Publicado: (2025)
por: Abnar, Samira, et al.
Publicado: (2025)
Cooperative Sense and Avoid for UAVs using Secondary Radar
por: Mohammadkarimi, Mostafa, et al.
Publicado: (2023)
por: Mohammadkarimi, Mostafa, et al.
Publicado: (2023)
On the Role of Initialization on the Implicit Bias in Deep Linear Networks
por: Gruber, Oria, et al.
Publicado: (2024)
por: Gruber, Oria, et al.
Publicado: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
por: Nakkiran, Preetum, et al.
Publicado: (2025)
por: Nakkiran, Preetum, et al.
Publicado: (2025)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
por: Devic, Siddartha, et al.
Publicado: (2025)
por: Devic, Siddartha, et al.
Publicado: (2025)
Normalizing Flows are Capable Generative Models
por: Zhai, Shuangfei, et al.
Publicado: (2024)
por: Zhai, Shuangfei, et al.
Publicado: (2024)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
por: Mallinar, Neil, et al.
Publicado: (2022)
por: Mallinar, Neil, et al.
Publicado: (2022)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
por: Aldeneh, Zakaria, et al.
Publicado: (2024)
Implicit Bias in Deep Linear Discriminant Analysis
por: Li, Jiawen
Publicado: (2026)
por: Li, Jiawen
Publicado: (2026)
A Formal Framework for Understanding Length Generalization in Transformers
por: Huang, Xinting, et al.
Publicado: (2024)
por: Huang, Xinting, et al.
Publicado: (2024)
How PARTs assemble into wholes: Learning the relative composition of images
por: Ayoughi, Melika, et al.
Publicado: (2025)
por: Ayoughi, Melika, et al.
Publicado: (2025)
How Far Can Transformers Reason? The Globality Barrier and Inductive Scratchpad
por: Abbe, Emmanuel, et al.
Publicado: (2024)
por: Abbe, Emmanuel, et al.
Publicado: (2024)
Path-Constrained Mixture-of-Experts
por: Gu, Zijin, et al.
Publicado: (2026)
por: Gu, Zijin, et al.
Publicado: (2026)
Neural Network Parameter-optimization of Gaussian pmDAGs
por: Saremi, Mehrzad
Publicado: (2023)
por: Saremi, Mehrzad
Publicado: (2023)
Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving
por: Wang, Linhan, et al.
Publicado: (2026)
por: Wang, Linhan, et al.
Publicado: (2026)
Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training
por: Sander, Tom, et al.
Publicado: (2024)
por: Sander, Tom, et al.
Publicado: (2024)
Optimal Implicit Bias in Linear Regression
por: Varma, Kanumuri Nithin, et al.
Publicado: (2025)
por: Varma, Kanumuri Nithin, et al.
Publicado: (2025)
Self-Distilled Depth Refinement with Noisy Poisson Fusion
por: Li, Jiaqi, et al.
Publicado: (2024)
por: Li, Jiaqi, et al.
Publicado: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
por: Cai, Yuhang, et al.
Publicado: (2025)
por: Cai, Yuhang, et al.
Publicado: (2025)
Foreign aid, conditionality, and ghost of the financing gap : a forgotten aspect of the aid debate / Thilak Ranaweera
por: Ranaweera, Thilak
Publicado: (2003)
por: Ranaweera, Thilak
Publicado: (2003)
Alternative paths to structural adjustment in Uzbekistan in a three-gap framework / Thilak Ranaweera
por: Ranaweera, Thilak
Publicado: (2003)
por: Ranaweera, Thilak
Publicado: (2003)
Market disequilibria and inflation in Uzbekistan, 1994-2000 / Thilak Ranaweera
por: Ranaweera, Thilak
Publicado: (2003)
por: Ranaweera, Thilak
Publicado: (2003)
How to Stay Curious while Avoiding Noisy TVs using Aleatoric Uncertainty Estimation
por: Mavor-Parker, Augustine N., et al.
Publicado: (2021)
por: Mavor-Parker, Augustine N., et al.
Publicado: (2021)
Obtenga capital sin riesgo / Asheesh Advani
por: Advani, Asheesh
Publicado: (2006)
por: Advani, Asheesh
Publicado: (2006)
When Small Models Are Right for Wrong Reasons: Process Verification for Trustworthy Agents
por: Advani, Laksh
Publicado: (2026)
por: Advani, Laksh
Publicado: (2026)
Trajectory Guard -- A Lightweight, Sequence-Aware Model for Real-Time Anomaly Detection in Agentic AI
por: Advani, Laksh
Publicado: (2026)
por: Advani, Laksh
Publicado: (2026)
Ejemplares similares
-
Vanishing Gradients in Reinforcement Finetuning of Language Models
por: Razin, Noam, et al.
Publicado: (2023) -
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
por: Huang, Chen, et al.
Publicado: (2026) -
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
por: Li, Xianhang, et al.
Publicado: (2025) -
Annotations Mitigate Post-Training Mode Collapse
por: Springer, Jacob Mitchell, et al.
Publicado: (2026) -
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
por: Littwin, Etai, et al.
Publicado: (2024)