Salvato in:
| Autori principali: | Springer, Jacob Mitchell, Advani, Madhu, Aichberger, Lukas, Bradley, Arwen, Malach, Eran, Saremi, Omid, Williamson, Sinead, Nakkiran, Preetum, Littwin, Etai, Raghunathan, Aditi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.09995 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
di: Littwin, Etai, et al.
Pubblicazione: (2024)
di: Littwin, Etai, et al.
Pubblicazione: (2024)
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
di: Malach, Eran, et al.
Pubblicazione: (2025)
di: Malach, Eran, et al.
Pubblicazione: (2025)
Vanishing Gradients in Reinforcement Finetuning of Language Models
di: Razin, Noam, et al.
Pubblicazione: (2023)
di: Razin, Noam, et al.
Pubblicazione: (2023)
Step-by-Step Diffusion: An Elementary Tutorial
di: Nakkiran, Preetum, et al.
Pubblicazione: (2024)
di: Nakkiran, Preetum, et al.
Pubblicazione: (2024)
Classifier-Free Guidance is a Predictor-Corrector
di: Bradley, Arwen, et al.
Pubblicazione: (2024)
di: Bradley, Arwen, et al.
Pubblicazione: (2024)
Trained on Tokens, Calibrated on Concepts: The Emergence of Semantic Calibration in LLMs
di: Nakkiran, Preetum, et al.
Pubblicazione: (2025)
di: Nakkiran, Preetum, et al.
Pubblicazione: (2025)
Trace Length is a Simple Uncertainty Signal in Reasoning Models
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
di: Devic, Siddartha, et al.
Pubblicazione: (2025)
Mechanisms of Projective Composition of Diffusion Models
di: Bradley, Arwen, et al.
Pubblicazione: (2025)
di: Bradley, Arwen, et al.
Pubblicazione: (2025)
Composition and Control with Distilled Energy Diffusion Models and Sequential Monte Carlo
di: Thornton, James, et al.
Pubblicazione: (2025)
di: Thornton, James, et al.
Pubblicazione: (2025)
When can transformers reason with abstract symbols?
di: Boix-Adsera, Enric, et al.
Pubblicazione: (2023)
di: Boix-Adsera, Enric, et al.
Pubblicazione: (2023)
Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers
di: Li, Xianhang, et al.
Pubblicazione: (2025)
di: Li, Xianhang, et al.
Pubblicazione: (2025)
Mitigating Bias in RAG: Controlling the Embedder
di: Kim, Taeyoun, et al.
Pubblicazione: (2025)
di: Kim, Taeyoun, et al.
Pubblicazione: (2025)
When is Multicalibration Post-Processing Necessary?
di: Hansen, Dutch, et al.
Pubblicazione: (2024)
di: Hansen, Dutch, et al.
Pubblicazione: (2024)
Enhancing JEPAs with Spatial Conditioning: Robust and Efficient Representation Learning
di: Littwin, Etai, et al.
Pubblicazione: (2024)
di: Littwin, Etai, et al.
Pubblicazione: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2024)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2024)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
di: Kotha, Suhas, et al.
Pubblicazione: (2023)
Auto-Regressive Next-Token Predictors are Universal Learners
di: Malach, Eran
Pubblicazione: (2023)
di: Malach, Eran
Pubblicazione: (2023)
Local Mechanisms of Compositional Generalization in Conditional Diffusion
di: Bradley, Arwen
Pubblicazione: (2025)
di: Bradley, Arwen
Pubblicazione: (2025)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
di: Huang, Chen, et al.
Pubblicazione: (2026)
di: Huang, Chen, et al.
Pubblicazione: (2026)
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity
di: Fu, Yicheng, et al.
Pubblicazione: (2024)
di: Fu, Yicheng, et al.
Pubblicazione: (2024)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
di: Karchmer, Ari, et al.
Pubblicazione: (2025)
di: Karchmer, Ari, et al.
Pubblicazione: (2025)
Uncertainty Quantification for LLM Function-Calling
di: Ye, Zihuiwen, et al.
Pubblicazione: (2026)
di: Ye, Zihuiwen, et al.
Pubblicazione: (2026)
Repetition Improves Language Model Embeddings
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2024)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2024)
Understanding the Influence of Synthetic Data for Text Embedders
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
di: Springer, Jacob Mitchell, et al.
Pubblicazione: (2025)
Early Data Exposure Improves Robustness to Subsequent Fine-Tuning
di: Feng, Lawrence, et al.
Pubblicazione: (2026)
di: Feng, Lawrence, et al.
Pubblicazione: (2026)
Unlocking the Working Memory of Large Language Models for Latent Reasoning
di: Aichberger, Lukas, et al.
Pubblicazione: (2026)
di: Aichberger, Lukas, et al.
Pubblicazione: (2026)
Distillation Scaling Laws
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
di: Busbridge, Dan, et al.
Pubblicazione: (2025)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
di: Mallinar, Neil, et al.
Pubblicazione: (2022)
di: Mallinar, Neil, et al.
Pubblicazione: (2022)
Mitigating Modal Imbalance in Multimodal Reasoning
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
Mode-Conditioning Unlocks Superior Test-Time Scaling
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
di: Wu, Chen Henry, et al.
Pubblicazione: (2025)
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
di: Zhao, Rosie, et al.
Pubblicazione: (2025)
di: Zhao, Rosie, et al.
Pubblicazione: (2025)
Rethinking Uncertainty Estimation in LLMs: A Principled Single-Sequence Measure
di: Aichberger, Lukas, et al.
Pubblicazione: (2024)
di: Aichberger, Lukas, et al.
Pubblicazione: (2024)
Differential Smoothing Mitigates Sharpening and Improves LLM Reasoning
di: Gai, Jingchu, et al.
Pubblicazione: (2025)
di: Gai, Jingchu, et al.
Pubblicazione: (2025)
A Taxonomy of Transcendence
di: Abreu, Natalie, et al.
Pubblicazione: (2025)
di: Abreu, Natalie, et al.
Pubblicazione: (2025)
LLM Priors for ERM over Programs
di: Singhal, Shivam, et al.
Pubblicazione: (2025)
di: Singhal, Shivam, et al.
Pubblicazione: (2025)
How Reinforcement Learning After Next-Token Prediction Facilitates Learning
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2025)
di: Tsilivis, Nikolaos, et al.
Pubblicazione: (2025)
On the Power of Decision Trees in Auto-Regressive Language Modeling
di: Gan, Yulu, et al.
Pubblicazione: (2024)
di: Gan, Yulu, et al.
Pubblicazione: (2024)
A Formal Framework for Understanding Length Generalization in Transformers
di: Huang, Xinting, et al.
Pubblicazione: (2024)
di: Huang, Xinting, et al.
Pubblicazione: (2024)
Self-Supervised Learning with Gaussian Processes
di: Duan, Yunshan, et al.
Pubblicazione: (2025)
di: Duan, Yunshan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
di: Littwin, Etai, et al.
Pubblicazione: (2024) -
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
di: Malach, Eran, et al.
Pubblicazione: (2025) -
Vanishing Gradients in Reinforcement Finetuning of Language Models
di: Razin, Noam, et al.
Pubblicazione: (2023) -
Step-by-Step Diffusion: An Elementary Tutorial
di: Nakkiran, Preetum, et al.
Pubblicazione: (2024) -
Classifier-Free Guidance is a Predictor-Corrector
di: Bradley, Arwen, et al.
Pubblicazione: (2024)