Balancing Label Quantity and Quality for Scalable Elicitation
Fuente:
arXiv
Saved in:
| Main Authors: | Mallen, Alex, Belrose, Nora |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023)
by: Mallen, Alex, et al.
Published: (2023)
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024)
by: Belrose, Nora, et al.
Published: (2024)
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)
by: Scherlis, Adam, et al.
Published: (2025)
Evaluating SAE interpretability without explanations
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Slowing Learning by Erasing Simple Features
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Sparse Autoencoders Trained on the Same Data Learn Different Features
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Converting MLPs into Polynomials in Closed Form
by: Belrose, Nora, et al.
Published: (2025)
by: Belrose, Nora, et al.
Published: (2025)
Partially Rewriting a Transformer in Natural Language
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Examining Two Hop Reasoning Through Information Content Scaling
by: Johnston, David, et al.
Published: (2025)
by: Johnston, David, et al.
Published: (2025)
Eliciting Latent Predictions from Transformers with the Tuned Lens
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Binary Sparse Coding for Interpretability
by: Quirke, Lucia, et al.
Published: (2025)
by: Quirke, Lucia, et al.
Published: (2025)
Transcoders Beat Sparse Autoencoders for Interpretability
by: Paulo, Gonçalo, et al.
Published: (2025)
by: Paulo, Gonçalo, et al.
Published: (2025)
Refusal in LLMs is an Affine Function
by: Marshall, Thomas, et al.
Published: (2024)
by: Marshall, Thomas, et al.
Published: (2024)
Does Transformer Interpretability Transfer to RNNs?
by: Paulo, Gonçalo, et al.
Published: (2024)
by: Paulo, Gonçalo, et al.
Published: (2024)
Mechanistic Anomaly Detection for "Quirky" Language Models
by: Johnston, David O., et al.
Published: (2025)
by: Johnston, David O., et al.
Published: (2025)
Dual-Criterion Model Aggregation in Federated Learning: Balancing Data Quantity and Quality
by: Zhang, Haizhou, et al.
Published: (2024)
by: Zhang, Haizhou, et al.
Published: (2024)
Navigating Data Corruption in Machine Learning: Balancing Quality, Quantity, and Imputation Strategies
by: Liu, Qi, et al.
Published: (2024)
by: Liu, Qi, et al.
Published: (2024)
Don't Label Twice: Quantity Beats Quality when Comparing Binary Classifiers on a Budget
by: Dorner, Florian E., et al.
Published: (2024)
by: Dorner, Florian E., et al.
Published: (2024)
Mitigating Modality Quantity and Quality Imbalance in Multimodal Online Federated Learning
by: Wang, Heqiang, et al.
Published: (2025)
by: Wang, Heqiang, et al.
Published: (2025)
Scalable Label Distribution Learning for Multi-Label Classification
by: Zhao, Xingyu, et al.
Published: (2023)
by: Zhao, Xingyu, et al.
Published: (2023)
LEACE: Perfect linear concept erasure in closed form
by: Belrose, Nora, et al.
Published: (2023)
by: Belrose, Nora, et al.
Published: (2023)
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
by: Mallen, Alex, et al.
Published: (2024)
by: Mallen, Alex, et al.
Published: (2024)
Quality over Quantity: Demonstration Curation via Influence Functions for Data-Centric Robot Learning
by: Lee, Haeone, et al.
Published: (2026)
by: Lee, Haeone, et al.
Published: (2026)
Why Do Some Language Models Fake Alignment While Others Don't?
by: Sheshadri, Abhay, et al.
Published: (2025)
by: Sheshadri, Abhay, et al.
Published: (2025)
Beyond Quantities: Machine Learning-based Characterization of Inequality in Infrastructure Quality Provision in Cities
by: Li, Bo, et al.
Published: (2024)
by: Li, Bo, et al.
Published: (2024)
Quality over Quantity: Boosting Data Efficiency Through Ensembled Multimodal Data Curation
by: Xu, Jinda, et al.
Published: (2025)
by: Xu, Jinda, et al.
Published: (2025)
Noether's razor: Learning Conserved Quantities
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
by: van der Ouderaa, Tycho F. A., et al.
Published: (2024)
Balancing Label Imbalance in Federated Environments Using Only Mixup and Artificially-Labeled Noise
by: Sang, Kyle, et al.
Published: (2024)
by: Sang, Kyle, et al.
Published: (2024)
Is Training Data Quality or Quantity More Impactful to Small Language Model Performance?
by: Sajith, Aryan, et al.
Published: (2024)
by: Sajith, Aryan, et al.
Published: (2024)
Orthogonal Representation Learning for Estimating Causal Quantities
by: Melnychuk, Valentyn, et al.
Published: (2025)
by: Melnychuk, Valentyn, et al.
Published: (2025)
ActiveCQ: Active Estimation of Causal Quantities
by: Gao, Erdun, et al.
Published: (2025)
by: Gao, Erdun, et al.
Published: (2025)
The Elicitation Game: Evaluating Capability Elicitation Techniques
by: Hofstätter, Felix, et al.
Published: (2025)
by: Hofstätter, Felix, et al.
Published: (2025)
From Overfitting to Robustness: Quantity, Quality, and Variety Oriented Negative Sample Selection in Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2024)
by: Ali, Adnan, et al.
Published: (2024)
Quality over Quantity: An Effective Large-Scale Data Reduction Strategy Based on Pointwise V-Information
by: Chen, Fei, et al.
Published: (2025)
by: Chen, Fei, et al.
Published: (2025)
Data-Efficient Contrastive Language-Image Pretraining: Prioritizing Data Quality over Quantity
by: Joshi, Siddharth, et al.
Published: (2024)
by: Joshi, Siddharth, et al.
Published: (2024)
Zero-to-Strong Generalization: Eliciting Strong Capabilities of Large Language Models Iteratively without Gold Labels
by: Liu, Chaoqun, et al.
Published: (2024)
by: Liu, Chaoqun, et al.
Published: (2024)
DHIL-GT: Scalable Graph Transformer with Decoupled Hierarchy Labeling
by: Liao, Ningyi, et al.
Published: (2024)
by: Liao, Ningyi, et al.
Published: (2024)
Truthful Elicitation of Imprecise Forecasts
by: Singh, Anurag, et al.
Published: (2025)
by: Singh, Anurag, et al.
Published: (2025)
Similar Items
-
Eliciting Latent Knowledge from Quirky Language Models
by: Mallen, Alex, et al.
Published: (2023) -
Automatically Interpreting Millions of Features in Large Language Models
by: Paulo, Gonçalo, et al.
Published: (2024) -
Neural Networks Learn Statistics of Increasing Complexity
by: Belrose, Nora, et al.
Published: (2024) -
Understanding Gradient Descent through the Training Jacobian
by: Belrose, Nora, et al.
Published: (2024) -
Estimating the Probability of Sampling a Trained Neural Network at Random
by: Scherlis, Adam, et al.
Published: (2025)