Scaling Laws For Scalable Oversight
Fuente:
arXiv
Saved in:
| Main Authors: | Engels, Joshua, Baek, David D., Kantamneni, Subhash, Tegmark, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025)
by: Kantamneni, Subhash, et al.
Published: (2025)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024)
by: Kantamneni, Subhash, et al.
Published: (2024)
A Resource Model For Neural Scaling Law
by: Song, Jinyeop, et al.
Published: (2024)
by: Song, Jinyeop, et al.
Published: (2024)
Investigating Representation Universality: Case Study on Genealogical Representations
by: Baek, David D., et al.
Published: (2024)
by: Baek, David D., et al.
Published: (2024)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
by: Li, Yuxiao, et al.
Published: (2024)
by: Li, Yuxiao, et al.
Published: (2024)
A Neural Scaling Law from Lottery Ticket Ensembling
by: Liu, Ziming, et al.
Published: (2023)
by: Liu, Ziming, et al.
Published: (2023)
Dense SAE Latents Are Features, Not Bugs
by: Sun, Xiaoqing, et al.
Published: (2025)
by: Sun, Xiaoqing, et al.
Published: (2025)
Language Models Represent Space and Time
by: Gurnee, Wes, et al.
Published: (2023)
by: Gurnee, Wes, et al.
Published: (2023)
Scaling Legal AI: Benchmarking Mamba and Transformers for Statutory Classification and Case Law Retrieval
by: Maurya, Anuraj
Published: (2025)
by: Maurya, Anuraj
Published: (2025)
Scaling Laws Do Not Scale
by: Diaz, Fernando, et al.
Published: (2023)
by: Diaz, Fernando, et al.
Published: (2023)
Generative AI Training and Copyright Law
by: Stober, Sebastian, et al.
Published: (2025)
by: Stober, Sebastian, et al.
Published: (2025)
Limits of trust in medical AI
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
Are clinicians ethically obligated to disclose their use of medical machine learning systems to patients?
by: Hatherley, Joshua
Published: (2025)
by: Hatherley, Joshua
Published: (2025)
Towards Scalable Oversight via Partitioned Human Supervision
by: Yin, Ren, et al.
Published: (2025)
by: Yin, Ren, et al.
Published: (2025)
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models
by: Potham, Ram, et al.
Published: (2025)
by: Potham, Ram, et al.
Published: (2025)
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
by: Wen, Xueru, et al.
Published: (2025)
by: Wen, Xueru, et al.
Published: (2025)
Are There Exceptions to Goodhart's Law? On the Moral Justification of Fairness-Aware Machine Learning
by: Weerts, Hilde, et al.
Published: (2022)
by: Weerts, Hilde, et al.
Published: (2022)
Training Foundation Models as Data Compression: On Information, Model Weights and Copyright Law
by: Franceschelli, Giorgio, et al.
Published: (2024)
by: Franceschelli, Giorgio, et al.
Published: (2024)
Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses
by: Cahyono, Joshua Adrian, et al.
Published: (2025)
by: Cahyono, Joshua Adrian, et al.
Published: (2025)
Transformers in Healthcare: A Survey
by: Nerella, Subhash, et al.
Published: (2023)
by: Nerella, Subhash, et al.
Published: (2023)
The Remarkable Robustness of LLMs: Stages of Inference?
by: Lad, Vedang, et al.
Published: (2024)
by: Lad, Vedang, et al.
Published: (2024)
Do Two AI Scientists Agree?
by: Fu, Xinghong, et al.
Published: (2025)
by: Fu, Xinghong, et al.
Published: (2025)
Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment
by: Sargeant, Holli, et al.
Published: (2024)
by: Sargeant, Holli, et al.
Published: (2024)
Neural Thermodynamic Laws for Large Language Model Training
by: Liu, Ziming, et al.
Published: (2025)
by: Liu, Ziming, et al.
Published: (2025)
Position: Model Collapse Does Not Mean What You Think
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
A New Paradigm for Counterfactual Reasoning in Fairness and Recourse
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Steering LLMs via Scalable Interactive Oversight
by: Zhou, Enyu, et al.
Published: (2026)
by: Zhou, Enyu, et al.
Published: (2026)
Persona-Based Simulation of Human Opinion at Population Scale
by: Li, Mao, et al.
Published: (2026)
by: Li, Mao, et al.
Published: (2026)
CarbonScaling: Extending Neural Scaling Laws for Carbon Footprint in Large Language Models
by: Jiang, Lei, et al.
Published: (2025)
by: Jiang, Lei, et al.
Published: (2025)
Low-Rank Adapting Models for Sparse Autoencoders
by: Chen, Matthew, et al.
Published: (2025)
by: Chen, Matthew, et al.
Published: (2025)
Decomposing The Dark Matter of Sparse Autoencoders
by: Engels, Joshua, et al.
Published: (2024)
by: Engels, Joshua, et al.
Published: (2024)
Towards AI-$45^{\circ}$ Law: A Roadmap to Trustworthy AGI
by: Yang, Chao, et al.
Published: (2024)
by: Yang, Chao, et al.
Published: (2024)
NYU CTF Bench: A Scalable Open-Source Benchmark Dataset for Evaluating LLMs in Offensive Security
by: Shao, Minghao, et al.
Published: (2024)
by: Shao, Minghao, et al.
Published: (2024)
Mastery Guided Non-parametric Clustering to Scale-up Strategy Prediction
by: Shakya, Anup, et al.
Published: (2024)
by: Shakya, Anup, et al.
Published: (2024)
Between Randomness and Arbitrariness: Some Lessons for Reliable Machine Learning at Scale
by: Cooper, A. Feder
Published: (2024)
by: Cooper, A. Feder
Published: (2024)
AI-powered Digital Framework for Personalized Economical Quality Learning at Scale
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
by: VatandoustMohammadieh, Mrzieh, et al.
Published: (2024)
Questionnaire Responses Do not Capture the Safety of AI Agents
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
by: Hellrigel-Holderbaum, Max, et al.
Published: (2026)
SUV: Scalable Large Language Model Copyright Compliance with Regularized Selective Unlearning
by: Xu, Tianyang, et al.
Published: (2025)
by: Xu, Tianyang, et al.
Published: (2025)
Similar Items
-
Language Models Use Trigonometry to Do Addition
by: Kantamneni, Subhash, et al.
Published: (2025) -
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
by: Kantamneni, Subhash, et al.
Published: (2025) -
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
by: Kantamneni, Subhash, et al.
Published: (2024) -
OptPDE: Discovering Novel Integrable Systems via AI-Human Collaboration
by: Kantamneni, Subhash, et al.
Published: (2024) -
A Resource Model For Neural Scaling Law
by: Song, Jinyeop, et al.
Published: (2024)