Beyond the Loss Curve: Scaling Laws, Active Learning, and the Limits of Learning from Exact Posteriors
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khorasani, Arian, Chen, Nathaniel, Oswal, Yug D, Gopalan, Akshat Santhana, Kolemen, Egemen, Shwartz-Ziv, Ravid |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
von: Patel, Niket, et al.
Veröffentlicht: (2024)
von: Patel, Niket, et al.
Veröffentlicht: (2024)
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
von: Shani, Chen, et al.
Veröffentlicht: (2025)
von: Shani, Chen, et al.
Veröffentlicht: (2025)
Video Representation Learning with Joint-Embedding Predictive Architectures
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024)
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2023)
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2023)
A Significantly Better Class of Activation Functions Than ReLU Like Activation Functions
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2024)
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2024)
Bending the Scaling Law Curve in Large-Scale Recommendation Systems
von: Ding, Qin, et al.
Veröffentlicht: (2026)
von: Ding, Qin, et al.
Veröffentlicht: (2026)
Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs
von: Chen, Angelica, et al.
Veröffentlicht: (2023)
von: Chen, Angelica, et al.
Veröffentlicht: (2023)
Cross-Temporal Attention Fusion (CTAF) for Multimodal Physiological Signals in Self-Supervised Learning
von: Khorasani, Arian, et al.
Veröffentlicht: (2026)
von: Khorasani, Arian, et al.
Veröffentlicht: (2026)
Variance-Covariance Regularization Improves Representation Learning
von: Zhu, Jiachen, et al.
Veröffentlicht: (2023)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2023)
Antislop: A Comprehensive Framework for Identifying and Eliminating Repetitive Patterns in Language Models
von: Paech, Samuel, et al.
Veröffentlicht: (2025)
von: Paech, Samuel, et al.
Veröffentlicht: (2025)
Counter-Geoengineering: Feasibility and Policy Implications for a Geoengineered World
von: de Bolle, Felipe, et al.
Veröffentlicht: (2024)
von: de Bolle, Felipe, et al.
Veröffentlicht: (2024)
Efficient Vectorized Backpropagation Algorithms for Training Feedforward Networks Composed of Quadratic Neurons
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2023)
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2023)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
AI Must Embrace Specialization via Superhuman Adaptable Intelligence
von: Goldfeder, Judah, et al.
Veröffentlicht: (2026)
von: Goldfeder, Judah, et al.
Veröffentlicht: (2026)
The Entropy Enigma: Success and Failure of Entropy Minimization
von: Press, Ori, et al.
Veröffentlicht: (2024)
von: Press, Ori, et al.
Veröffentlicht: (2024)
Alternate Loss Functions for Classification and Robust Regression Can Improve the Accuracy of Artificial Neural Networks
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2023)
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2023)
Does Representation Matter? Exploring Intermediate Layers in Large Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2024)
von: Skean, Oscar, et al.
Veröffentlicht: (2024)
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
von: Shaar, Eitan, et al.
Veröffentlicht: (2026)
A Mutual Information Lower Bound for Multimodal Regression Active Learning
von: Guilhoto, Leonardo Ferreira, et al.
Veröffentlicht: (2026)
von: Guilhoto, Leonardo Ferreira, et al.
Veröffentlicht: (2026)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2025)
On Training in Imagination
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
von: Timor, Nadav, et al.
Veröffentlicht: (2026)
Model Editing at Scale leads to Gradual and Catastrophic Forgetting
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
von: Gupta, Akshat, et al.
Veröffentlicht: (2024)
Fast Power Curve Approximation for Posterior Analyses
von: Hagar, Luke, et al.
Veröffentlicht: (2023)
von: Hagar, Luke, et al.
Veröffentlicht: (2023)
Information-Theoretic Limits on Exact Subgraph Alignment Problem
von: Shiu, Chun Hei Michael, et al.
Veröffentlicht: (2026)
von: Shiu, Chun Hei Michael, et al.
Veröffentlicht: (2026)
Limited Improvement of Connectivity in Scale-Free Networks by Increasing the Power-Law Exponent
von: Mou, Yingzhou, et al.
Veröffentlicht: (2025)
von: Mou, Yingzhou, et al.
Veröffentlicht: (2025)
The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
von: Janiak, Denis, et al.
Veröffentlicht: (2025)
Neural Polar Decoders for DNA Data Storage
von: Aharoni, Ziv, et al.
Veröffentlicht: (2025)
von: Aharoni, Ziv, et al.
Veröffentlicht: (2025)
Pufferfish Privacy: An Information-Theoretic Study
von: Nuradha, Theshani, et al.
Veröffentlicht: (2022)
von: Nuradha, Theshani, et al.
Veröffentlicht: (2022)
Soft Clustering Anchors for Self-Supervised Speech Representation Learning in Joint Embedding Prediction Architectures
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
von: Ioannides, Georgios, et al.
Veröffentlicht: (2026)
Age of Information Optimization with Preemption Strategies for Correlated Systems
von: Erbayat, Egemen, et al.
Veröffentlicht: (2025)
von: Erbayat, Egemen, et al.
Veröffentlicht: (2025)
Age of Information Optimization and State Error Analysis for Correlated Multi-Process Multi-Sensor Systems
von: Erbayat, Egemen, et al.
Veröffentlicht: (2024)
von: Erbayat, Egemen, et al.
Veröffentlicht: (2024)
Age of Information Optimization in Distributed Sensor Networks with Half-Duplex Channels
von: Zou, Peng, et al.
Veröffentlicht: (2026)
von: Zou, Peng, et al.
Veröffentlicht: (2026)
Neural Polar Decoders for Deletion Channels
von: Aharoni, Ziv, et al.
Veröffentlicht: (2025)
von: Aharoni, Ziv, et al.
Veröffentlicht: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
Team ACK at SemEval-2025 Task 2: Beyond Word-for-Word Machine Translation for English-Korean Pairs
von: Lee, Daniel, et al.
Veröffentlicht: (2025)
von: Lee, Daniel, et al.
Veröffentlicht: (2025)
Scaling Laws for Cross-Encoder Reranking
von: Seetharaman, Rahul, et al.
Veröffentlicht: (2026)
von: Seetharaman, Rahul, et al.
Veröffentlicht: (2026)
Leave-One-Out Learning with Log-Loss
von: Fogel, Yaniv, et al.
Veröffentlicht: (2025)
von: Fogel, Yaniv, et al.
Veröffentlicht: (2025)
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs
von: Nguyen, Minh Nhat, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh Nhat, et al.
Veröffentlicht: (2024)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
von: LeVi, Amit, et al.
Veröffentlicht: (2025)
von: LeVi, Amit, et al.
Veröffentlicht: (2025)
Rate-In: Information-Driven Adaptive Dropout Rates for Improved Inference-Time Uncertainty Estimation
von: Zeevi, Tal, et al.
Veröffentlicht: (2024)
von: Zeevi, Tal, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning to Compress: Local Rank and Information Compression in Deep Neural Networks
von: Patel, Niket, et al.
Veröffentlicht: (2024) -
From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning
von: Shani, Chen, et al.
Veröffentlicht: (2025) -
Video Representation Learning with Joint-Embedding Predictive Architectures
von: Drozdov, Katrina, et al.
Veröffentlicht: (2024) -
An Information-Theoretic Perspective on Variance-Invariance-Covariance Regularization
von: Shwartz-Ziv, Ravid, et al.
Veröffentlicht: (2023) -
A Significantly Better Class of Activation Functions Than ReLU Like Activation Functions
von: Noel, Mathew Mithra, et al.
Veröffentlicht: (2024)