Non-Vacuous Generalization Bounds for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Lotfi, Sanae, Finzi, Marc, Kuang, Yilun, Rudner, Tim G. J., Goldblum, Micah, Wilson, Andrew Gordon |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
by: Lotfi, Sanae, et al.
Published: (2024)
by: Lotfi, Sanae, et al.
Published: (2024)
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025)
by: Marek, Martin, et al.
Published: (2025)
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
by: Goldblum, Micah, et al.
Published: (2023)
by: Goldblum, Micah, et al.
Published: (2023)
The Lie Derivative for Measuring Learned Equivariance
by: Gruver, Nate, et al.
Published: (2022)
by: Gruver, Nate, et al.
Published: (2022)
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)
by: Qiu, Shikai, et al.
Published: (2024)
Customizing the Inductive Biases of Softmax Attention using Structured Matrices
by: Kuang, Yilun, et al.
Published: (2025)
by: Kuang, Yilun, et al.
Published: (2025)
Large Language Models Are Zero-Shot Time Series Forecasters
by: Gruver, Nate, et al.
Published: (2023)
by: Gruver, Nate, et al.
Published: (2023)
Searching for Efficient Linear Layers over a Continuous Space of Structured Matrices
by: Potapczynski, Andres, et al.
Published: (2024)
by: Potapczynski, Andres, et al.
Published: (2024)
Learning Non-Vacuous Generalization Bounds from Optimization
by: Tan, Chengli, et al.
Published: (2022)
by: Tan, Chengli, et al.
Published: (2022)
Non-Vacuous Generalization Bounds: Can Rescaling Invariances Help?
by: Rouchouse, Damien, et al.
Published: (2025)
by: Rouchouse, Damien, et al.
Published: (2025)
From Entropy to Epiplexity: Rethinking Information for Computationally Bounded Intelligence
by: Finzi, Marc, et al.
Published: (2026)
by: Finzi, Marc, et al.
Published: (2026)
A Study of Bayesian Neural Network Surrogates for Bayesian Optimization
by: Li, Yucen Lily, et al.
Published: (2023)
by: Li, Yucen Lily, et al.
Published: (2023)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
Model Merging is Secretly Certifiable: Non-Vacuous Generalisation Bounds for Low-Shot Learning
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
From Low Intrinsic Dimensionality to Non-Vacuous Generalization Bounds in Deep Multi-Task Learning
by: Zakerinia, Hossein, et al.
Published: (2025)
by: Zakerinia, Hossein, et al.
Published: (2025)
Knowing What You Know Is Not Enough: Large Language Model Confidences Don't Align With Their Actions
by: Pal, Arka, et al.
Published: (2025)
by: Pal, Arka, et al.
Published: (2025)
Large Language Models Must Be Taught to Know What They Don't Know
by: Kapoor, Sanyam, et al.
Published: (2024)
by: Kapoor, Sanyam, et al.
Published: (2024)
Mind the GAP: Improving Robustness to Subpopulation Shifts with Group-Aware Priors
by: Rudner, Tim G. J., et al.
Published: (2024)
by: Rudner, Tim G. J., et al.
Published: (2024)
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
by: Kuang, Yilun, et al.
Published: (2026)
by: Kuang, Yilun, et al.
Published: (2026)
Just How Flexible are Neural Networks in Practice?
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
by: Shwartz-Ziv, Ravid, et al.
Published: (2024)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Insertion Language Models: Sequence Generation with Arbitrary-Position Insertions
by: Patel, Dhruvesh, et al.
Published: (2025)
by: Patel, Dhruvesh, et al.
Published: (2025)
In-Context Clustering with Large Language Models
by: Wang, Ying, et al.
Published: (2025)
by: Wang, Ying, et al.
Published: (2025)
Radial-VCReg: More Informative Representation Learning Through Radial Gaussianization
by: Kuang, Yilun, et al.
Published: (2026)
by: Kuang, Yilun, et al.
Published: (2026)
Dynamic Delayed Tree Expansion For Improved Multi-Path Speculative Decoding
by: Thomas, Rahul, et al.
Published: (2026)
by: Thomas, Rahul, et al.
Published: (2026)
Out-of-Distribution Detection Methods Answer the Wrong Questions
by: Li, Yucen Lily, et al.
Published: (2025)
by: Li, Yucen Lily, et al.
Published: (2025)
Non-Vacuous Certification of Transport MCMC via Oscillation-Controlled Normalizing Flows
by: Hu, Jun
Published: (2026)
by: Hu, Jun
Published: (2026)
Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
by: Jain, Neel, et al.
Published: (2024)
by: Jain, Neel, et al.
Published: (2024)
On the Reliability of Watermarks for Large Language Models
by: Kirchenbauer, John, et al.
Published: (2023)
by: Kirchenbauer, John, et al.
Published: (2023)
Predicting the Performance of Black-box LLMs through Follow-up Queries
by: Sam, Dylan, et al.
Published: (2025)
by: Sam, Dylan, et al.
Published: (2025)
Improving Semantic Uncertainty Quantification in Language Model Question-Answering via Token-Level Temperature Scaling
by: Lamb, Tom A., et al.
Published: (2026)
by: Lamb, Tom A., et al.
Published: (2026)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
by: Souri, Hossein, et al.
Published: (2024)
by: Souri, Hossein, et al.
Published: (2024)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
by: Bhardwaj, Dhrupad, et al.
Published: (2025)
Bayesian Optimization of Antibodies Informed by a Generative Model of Evolving Sequences
by: Amin, Alan Nawzad, et al.
Published: (2024)
by: Amin, Alan Nawzad, et al.
Published: (2024)
vTune: Verifiable Fine-Tuning for LLMs Through Backdooring
by: Zhang, Eva, et al.
Published: (2024)
by: Zhang, Eva, et al.
Published: (2024)
Trade-offs in Ensembling, Merging and Routing Among Parameter-Efficient Experts
by: Lotfi, Sanae, et al.
Published: (2026)
by: Lotfi, Sanae, et al.
Published: (2026)
Diffusing Differentiable Representations
by: Savani, Yash, et al.
Published: (2024)
by: Savani, Yash, et al.
Published: (2024)
Privacy-Preserving Mechanisms Enable Cheap Verifiable Inference of LLMs
by: Pal, Arka, et al.
Published: (2026)
by: Pal, Arka, et al.
Published: (2026)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
by: Chen, Haoran, et al.
Published: (2024)
by: Chen, Haoran, et al.
Published: (2024)
Similar Items
-
Unlocking Tokens as Data Points for Generalization Bounds on Larger Language Models
by: Lotfi, Sanae, et al.
Published: (2024) -
Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation Is Wasteful
by: Marek, Martin, et al.
Published: (2025) -
The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning
by: Goldblum, Micah, et al.
Published: (2023) -
The Lie Derivative for Measuring Learned Equivariance
by: Gruver, Nate, et al.
Published: (2022) -
Compute Better Spent: Replacing Dense Layers with Structured Matrices
by: Qiu, Shikai, et al.
Published: (2024)