Scales++: Compute Efficient Evaluation Subset Selection with Cognitive Scales Embeddings
Fuente:
arXiv
Saved in:
| Main Authors: | Bean, Andrew M., Seedat, Nabeel, Chen, Shengzhuang, Schwarz, Jonathan Richard |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands
by: Yu, Fangyi, et al.
Published: (2026)
by: Yu, Fangyi, et al.
Published: (2026)
ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
by: Chen, Shengzhuang, et al.
Published: (2025)
by: Chen, Shengzhuang, et al.
Published: (2025)
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023)
by: Sun, Hao, et al.
Published: (2023)
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
by: Seedat, Nabeel, et al.
Published: (2024)
by: Seedat, Nabeel, et al.
Published: (2024)
Large Language Models to Enhance Bayesian Optimization
by: Liu, Tennison, et al.
Published: (2024)
by: Liu, Tennison, et al.
Published: (2024)
Self-Healing Machine Learning: A Framework for Autonomous Adaptation in Real-World Environments
by: Rauba, Paulius, et al.
Published: (2024)
by: Rauba, Paulius, et al.
Published: (2024)
DAGnosis: Localized Identification of Data Inconsistencies using Structures
by: Huynh, Nicolas, et al.
Published: (2024)
by: Huynh, Nicolas, et al.
Published: (2024)
Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimes
by: Seedat, Nabeel, et al.
Published: (2023)
by: Seedat, Nabeel, et al.
Published: (2023)
DC-Check: A Data-Centric AI checklist to guide the development of reliable machine learning systems
by: Seedat, Nabeel, et al.
Published: (2022)
by: Seedat, Nabeel, et al.
Published: (2022)
CoScale-RL: Efficient Post-Training by Co-Scaling Data and Computation
by: Chen, Yutong, et al.
Published: (2026)
by: Chen, Yutong, et al.
Published: (2026)
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
by: Yu, Fangyi, et al.
Published: (2025)
by: Yu, Fangyi, et al.
Published: (2025)
Automatic Expert Discovery in LLM Upcycling via Sparse Interpolated Mixture-of-Experts
by: Chen, Shengzhuang, et al.
Published: (2025)
by: Chen, Shengzhuang, et al.
Published: (2025)
PSNE: Efficient Spectral Sparsification Algorithms for Scaling Network Embedding
by: Lin, Longlong, et al.
Published: (2024)
by: Lin, Longlong, et al.
Published: (2024)
Influence-Guided Symbolic Regression: Scientific Discovery via LLM-Driven Equation Search with Granular Feedback
by: Saveliev, Evgeny S., et al.
Published: (2026)
by: Saveliev, Evgeny S., et al.
Published: (2026)
DataS^3: Dataset Subset Selection for Specialization
by: Hulkund, Neha, et al.
Published: (2025)
by: Hulkund, Neha, et al.
Published: (2025)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Bandit Guided Submodular Curriculum for Adaptive Subset Selection
by: Chanda, Prateek, et al.
Published: (2025)
by: Chanda, Prateek, et al.
Published: (2025)
You Only Train Once: Differentiable Subset Selection for Omics Data
by: Chopard, Daphné, et al.
Published: (2025)
by: Chopard, Daphné, et al.
Published: (2025)
Compute-Optimal LLMs Provably Generalize Better With Scale
by: Finzi, Marc, et al.
Published: (2025)
by: Finzi, Marc, et al.
Published: (2025)
NuTime: Numerically Multi-Scaled Embedding for Large-Scale Time-Series Pretraining
by: Lin, Chenguo, et al.
Published: (2023)
by: Lin, Chenguo, et al.
Published: (2023)
Selecting Subsets of Source Data for Transfer Learning with Applications in Metal Additive Manufacturing
by: Tang, Yifan, et al.
Published: (2024)
by: Tang, Yifan, et al.
Published: (2024)
Efficient Exploration at Scale
by: Asghari, Seyed Mohammad, et al.
Published: (2026)
by: Asghari, Seyed Mohammad, et al.
Published: (2026)
Scaling Embeddings Outperforms Scaling Experts in Language Models
by: Liu, Hong, et al.
Published: (2026)
by: Liu, Hong, et al.
Published: (2026)
LinkedIn Post Embeddings: Industrial Scale Embedding Generation and Usage across LinkedIn
by: Ramanujam, Sudarshan Srinivasa, et al.
Published: (2024)
by: Ramanujam, Sudarshan Srinivasa, et al.
Published: (2024)
Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise
by: Shubham, Kumar, et al.
Published: (2026)
by: Shubham, Kumar, et al.
Published: (2026)
Market-Driven Subset Selection for Budgeted Training
by: Jha, Ashish, et al.
Published: (2025)
by: Jha, Ashish, et al.
Published: (2025)
CLIP-Embed-KD: Computationally Efficient Knowledge Distillation Using Embeddings as Teachers
by: Nair, Lakshmi
Published: (2024)
by: Nair, Lakshmi
Published: (2024)
Class-based Subset Selection for Transfer Learning under Extreme Label Shift
by: Goyal, Akul, et al.
Published: (2024)
by: Goyal, Akul, et al.
Published: (2024)
Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization
by: Shen, Qianli, et al.
Published: (2024)
by: Shen, Qianli, et al.
Published: (2024)
The Art of Scaling Reinforcement Learning Compute for LLMs
by: Khatri, Devvrit, et al.
Published: (2025)
by: Khatri, Devvrit, et al.
Published: (2025)
Evaluation-driven Scaling for Scientific Discovery
by: Ye, Haotian, et al.
Published: (2026)
by: Ye, Haotian, et al.
Published: (2026)
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
by: Cheng, Zhoujun, et al.
Published: (2026)
by: Cheng, Zhoujun, et al.
Published: (2026)
CauScale: Neural Causal Discovery at Scale
by: Peng, Bo, et al.
Published: (2026)
by: Peng, Bo, et al.
Published: (2026)
Scaling Strategy, Not Compute: A Stand-Alone, Open-Source StarCraft II Benchmark for Accessible Reinforcement Learning Research
by: Panda, Sourav, et al.
Published: (2026)
by: Panda, Sourav, et al.
Published: (2026)
Efficient Deep Learning Infrastructures for Embedded Computing Systems: A Comprehensive Survey and Future Envision
by: Luo, Xiangzhong, et al.
Published: (2024)
by: Luo, Xiangzhong, et al.
Published: (2024)
LOCUS: Low-Dimensional Model Embeddings for Efficient Model Exploration, Comparison, and Selection
by: Patel, Shivam, et al.
Published: (2026)
by: Patel, Shivam, et al.
Published: (2026)
Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
by: Brown, Bradley, et al.
Published: (2024)
by: Brown, Bradley, et al.
Published: (2024)
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
by: Huml, JR, et al.
Published: (2026)
by: Huml, JR, et al.
Published: (2026)
Scaling Laws for Data-Efficient Visual Transfer Learning
by: Yang, Wenxuan, et al.
Published: (2025)
by: Yang, Wenxuan, et al.
Published: (2025)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
Similar Items
-
To Whom Do Language Models Align? Measuring Principal Hierarchies Under High-Stakes Competing Demands
by: Yu, Fangyi, et al.
Published: (2026) -
ADMIRE-BayesOpt: Accelerated Data MIxture RE-weighting for Language Models with Bayesian Optimization
by: Chen, Shengzhuang, et al.
Published: (2025) -
When is Off-Policy Evaluation (Reward Modeling) Useful in Contextual Bandits? A Data-Centric Perspective
by: Sun, Hao, et al.
Published: (2023) -
You can't handle the (dirty) truth: Data-centric insights improve pseudo-labeling
by: Seedat, Nabeel, et al.
Published: (2024) -
Large Language Models to Enhance Bayesian Optimization
by: Liu, Tennison, et al.
Published: (2024)