From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
Fuente:
arXiv
Guardado en:
| Autores principales: | Finzi, Marc, Qiu, Shikai, Jiang, Yiding, Izmailov, Pavel, Kolter, J. Zico, Wilson, Andrew Gordon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Large Language Models Are Zero-Shot Time Series Forecasters
por: Gruver, Nate, et al.
Publicado: (2023)
por: Gruver, Nate, et al.
Publicado: (2023)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
por: Qiu, Shikai, et al.
Publicado: (2024)
por: Qiu, Shikai, et al.
Publicado: (2024)
Predicting the Performance of Black-box LLMs through Follow-up Queries
por: Sam, Dylan, et al.
Publicado: (2025)
por: Sam, Dylan, et al.
Publicado: (2025)
Diffusing Differentiable Representations
por: Savani, Yash, et al.
Publicado: (2024)
por: Savani, Yash, et al.
Publicado: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
por: Finzi, Marc, et al.
Publicado: (2025)
por: Finzi, Marc, et al.
Publicado: (2025)
Forgetting in Language Models: Capacity, Optimization, and Self-Generated Replay
por: Marek, Martin, et al.
Publicado: (2026)
por: Marek, Martin, et al.
Publicado: (2026)
Looking beyond the next token
por: Thankaraj, Abitha, et al.
Publicado: (2025)
por: Thankaraj, Abitha, et al.
Publicado: (2025)
Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws
por: Jiang, Yiding, et al.
Publicado: (2024)
por: Jiang, Yiding, et al.
Publicado: (2024)
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
por: Lotfi, Sanae, et al.
Publicado: (2024)
por: Lotfi, Sanae, et al.
Publicado: (2024)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
por: Goldblum, Micah, et al.
Publicado: (2023)
por: Goldblum, Micah, et al.
Publicado: (2023)
Non-Vacuous Generalization Bounds for Large Language Models
por: Lotfi, Sanae, et al.
Publicado: (2023)
por: Lotfi, Sanae, et al.
Publicado: (2023)
Rethinking Distance Metrics for Counterfactual Explainability
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
por: Potapczynski, Andres, et al.
Publicado: (2024)
por: Potapczynski, Andres, et al.
Publicado: (2024)
Scaling Collapse Reveals Universal Dynamics in Compute-Optimally Trained Neural Networks
por: Qiu, Shikai, et al.
Publicado: (2025)
por: Qiu, Shikai, et al.
Publicado: (2025)
The Lie Derivative for Measuring Learned Equivariance
por: Gruver, Nate, et al.
Publicado: (2022)
por: Gruver, Nate, et al.
Publicado: (2022)
AcceleratedLiNGAM: Learning Causal DAGs at the speed of GPUs
por: Akinwande, Victor, et al.
Publicado: (2024)
por: Akinwande, Victor, et al.
Publicado: (2024)
FUSE-ing Language Models: Zero-Shot Adapter Discovery for Prompt Optimization Across Tokenizers
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
por: Williams, Joshua Nathaniel, et al.
Publicado: (2024)
Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
por: Huang, Benhao, et al.
Publicado: (2026)
por: Huang, Benhao, et al.
Publicado: (2026)
Why is SAM Robust to Label Noise?
por: Baek, Christina, et al.
Publicado: (2024)
por: Baek, Christina, et al.
Publicado: (2024)
Rethinking LLM Memorization through the Lens of Adversarial Compression
por: Schwarzschild, Avi, et al.
Publicado: (2024)
por: Schwarzschild, Avi, et al.
Publicado: (2024)
Mimetic Initialization of MLPs
por: Trockman, Asher, et al.
Publicado: (2026)
por: Trockman, Asher, et al.
Publicado: (2026)
Provably Bounding Neural Network Preimages
por: Kotha, Suhas, et al.
Publicado: (2023)
por: Kotha, Suhas, et al.
Publicado: (2023)
Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across Scales
por: Qiu, Shikai, et al.
Publicado: (2025)
por: Qiu, Shikai, et al.
Publicado: (2025)
Evaluating Language Model Reasoning about Confidential Information
por: Sam, Dylan, et al.
Publicado: (2025)
por: Sam, Dylan, et al.
Publicado: (2025)
Scaling Laws for Data Filtering -- Data Curation cannot be Compute Agnostic
por: Goyal, Sachin, et al.
Publicado: (2024)
por: Goyal, Sachin, et al.
Publicado: (2024)
Training a Generally Curious Agent
por: Tajwar, Fahim, et al.
Publicado: (2025)
por: Tajwar, Fahim, et al.
Publicado: (2025)
One-Step Diffusion Distillation via Deep Equilibrium Models
por: Geng, Zhengyang, et al.
Publicado: (2023)
por: Geng, Zhengyang, et al.
Publicado: (2023)
Measuring Five-Nines Reliability: Sample-Efficient LLM Evaluation in Saturated Benchmarks
por: Kim, Eungyeup, et al.
Publicado: (2026)
por: Kim, Eungyeup, et al.
Publicado: (2026)
Generative Posterior Networks for Approximately Bayesian Epistemic Uncertainty Estimation
por: Roderick, Melrose, et al.
Publicado: (2023)
por: Roderick, Melrose, et al.
Publicado: (2023)
Neural Network Verification with Branch-and-Bound for General Nonlinearities
por: Shi, Zhouxing, et al.
Publicado: (2024)
por: Shi, Zhouxing, et al.
Publicado: (2024)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
por: Kuang, Yilun, et al.
Publicado: (2025)
por: Kuang, Yilun, et al.
Publicado: (2025)
Can a Confident Prior Replace a Cold Posterior?
por: Marek, Martin, et al.
Publicado: (2024)
por: Marek, Martin, et al.
Publicado: (2024)
An Axiomatic Approach to Model-Agnostic Concept Explanations
por: Feng, Zhili, et al.
Publicado: (2024)
por: Feng, Zhili, et al.
Publicado: (2024)
The Mixing method: low-rank coordinate descent for semidefinite programming with diagonal constraints
por: Wang, Po-Wei, et al.
Publicado: (2017)
por: Wang, Po-Wei, et al.
Publicado: (2017)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
por: Goyal, Sachin, et al.
Publicado: (2024)
por: Goyal, Sachin, et al.
Publicado: (2024)
Massive Activations in Large Language Models
por: Sun, Mingjie, et al.
Publicado: (2024)
por: Sun, Mingjie, et al.
Publicado: (2024)
From Variance to Veracity: Unbundling and Mitigating Gradient Variance in Differentiable Bundle Adjustment Layers
por: Gurumurthy, Swaminathan, et al.
Publicado: (2024)
por: Gurumurthy, Swaminathan, et al.
Publicado: (2024)
Transferring Knowledge from Large Foundation Models to Small Downstream Models
por: Qiu, Shikai, et al.
Publicado: (2024)
por: Qiu, Shikai, et al.
Publicado: (2024)
ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use
por: Tien, Jeremy, et al.
Publicado: (2026)
por: Tien, Jeremy, et al.
Publicado: (2026)
Safety Pretraining: Toward the Next Generation of Safe AI
por: Maini, Pratyush, et al.
Publicado: (2025)
por: Maini, Pratyush, et al.
Publicado: (2025)
Ejemplares similares
-
Large Language Models Are Zero-Shot Time Series Forecasters
por: Gruver, Nate, et al.
Publicado: (2023) -
Compute Better Spent: Replacing Dense Layers with Structured Matrices
por: Qiu, Shikai, et al.
Publicado: (2024) -
Predicting the Performance of Black-box LLMs through Follow-up Queries
por: Sam, Dylan, et al.
Publicado: (2025) -
Diffusing Differentiable Representations
por: Savani, Yash, et al.
Publicado: (2024) -
Compute-Optimal LLMs Provably Generalize Better With Scale
por: Finzi, Marc, et al.
Publicado: (2025)