DsDm: Model-Aware Dataset Selection with Datamodels
Fuente:
arXiv
Guardado en:
| Autores principales: | Engstrom, Logan, Feldmann, Axel, Madry, Aleksander |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
por: Dass, Shivin, et al.
Publicado: (2025)
por: Dass, Shivin, et al.
Publicado: (2025)
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
por: Khaddaj, Alaa, et al.
Publicado: (2025)
por: Khaddaj, Alaa, et al.
Publicado: (2025)
Optimizing ML Training with Metagradient Descent
por: Engstrom, Logan, et al.
Publicado: (2025)
por: Engstrom, Logan, et al.
Publicado: (2025)
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
por: Jain, Saachi, et al.
Publicado: (2024)
por: Jain, Saachi, et al.
Publicado: (2024)
Attribute-to-Delete: Machine Unlearning via Datamodel Matching
por: Georgiev, Kristian, et al.
Publicado: (2024)
por: Georgiev, Kristian, et al.
Publicado: (2024)
Decomposing and Editing Predictions by Modeling Model Computation
por: Shah, Harshay, et al.
Publicado: (2024)
por: Shah, Harshay, et al.
Publicado: (2024)
Ask Your Distribution Shift if Pre-Training is Right for You
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
Do Large Language Model Benchmarks Test Reliability?
por: Vendrow, Joshua, et al.
Publicado: (2025)
por: Vendrow, Joshua, et al.
Publicado: (2025)
ContextCite: Attributing Model Generation to Context
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2024)
Distilled Datamodel with Reverse Gradient Matching
por: Ye, Jingwen, et al.
Publicado: (2024)
por: Ye, Jingwen, et al.
Publicado: (2024)
Learning to Attribute with Attention
por: Cohen-Wang, Benjamin, et al.
Publicado: (2025)
por: Cohen-Wang, Benjamin, et al.
Publicado: (2025)
User Strategization and Trustworthy Algorithms
por: Cen, Sarah H., et al.
Publicado: (2023)
por: Cen, Sarah H., et al.
Publicado: (2023)
MAGIC: Near-Optimal Data Attribution for Deep Learning
por: Ilyas, Andrew, et al.
Publicado: (2025)
por: Ilyas, Andrew, et al.
Publicado: (2025)
DmC: Nearest Neighbor Guidance Diffusion Model for Offline Cross-domain Reinforcement Learning
por: Van, Linh Le Pham, et al.
Publicado: (2025)
por: Van, Linh Le Pham, et al.
Publicado: (2025)
Large-Scale, Longitudinal Study of Large Language Models During the 2024 US Election Season
por: Cen, Sarah H., et al.
Publicado: (2025)
por: Cen, Sarah H., et al.
Publicado: (2025)
AI Supply Chains: An Emerging Ecosystem of AI Actors, Products, and Services
por: Hopkins, Aspen, et al.
Publicado: (2025)
por: Hopkins, Aspen, et al.
Publicado: (2025)
Measuring Strategization in Recommendation: Users Adapt Their Behavior to Shape Future Content
por: Cen, Sarah H., et al.
Publicado: (2024)
por: Cen, Sarah H., et al.
Publicado: (2024)
Decision-Aware Proximal Bridge Learning for Optimal Treatment Selection
por: Garriga, Tomàs, et al.
Publicado: (2026)
por: Garriga, Tomàs, et al.
Publicado: (2026)
Layer-wise Lipschitz-Product Control for Deep Kolmogorov--Arnold Network Representations of Compositionally Structured Functions
por: Tankman, Aleksander
Publicado: (2026)
por: Tankman, Aleksander
Publicado: (2026)
Near-Infrared Hyperspectral Imaging Applications in Food Analysis -- Improving Algorithms and Methodologies
por: Engstrøm, Ole-Christian Galbo
Publicado: (2025)
por: Engstrøm, Ole-Christian Galbo
Publicado: (2025)
Deception Detection: From Static Texts to Multimodal Signals
por: Logan, Mandela
Publicado: (2025)
por: Logan, Mandela
Publicado: (2025)
Assessing the Potential of Masked Autoencoder Foundation Models in Predicting Downhole Metrics from Surface Drilling Data
por: Berezowski, Aleksander, et al.
Publicado: (2026)
por: Berezowski, Aleksander, et al.
Publicado: (2026)
GenSelect: A Generative Approach to Best-of-N
por: Toshniwal, Shubham, et al.
Publicado: (2025)
por: Toshniwal, Shubham, et al.
Publicado: (2025)
Decomposition of Small Transformer Models
por: Christensen, Casper L., et al.
Publicado: (2025)
por: Christensen, Casper L., et al.
Publicado: (2025)
Computation-Aware Gaussian Processes: Model Selection And Linear-Time Inference
por: Wenger, Jonathan, et al.
Publicado: (2024)
por: Wenger, Jonathan, et al.
Publicado: (2024)
Semantics-Aware Caching for Concept Learning
por: Teyou, Louis Mozart Kamdem, et al.
Publicado: (2026)
por: Teyou, Louis Mozart Kamdem, et al.
Publicado: (2026)
Distribution-Aware Feature Selection for SAEs
por: Oozeer, Narmeen, et al.
Publicado: (2025)
por: Oozeer, Narmeen, et al.
Publicado: (2025)
Decision-Aware Predictive Model Selection for Workforce Allocation
por: Stratman, Eric G., et al.
Publicado: (2024)
por: Stratman, Eric G., et al.
Publicado: (2024)
Fast Partition-Based Cross-Validation With Centering and Scaling for $\mathbf{X}^\mathbf{T}\mathbf{X}$ and $\mathbf{X}^\mathbf{T}\mathbf{Y}$
por: Engstrøm, Ole-Christian Galbo, et al.
Publicado: (2024)
por: Engstrøm, Ole-Christian Galbo, et al.
Publicado: (2024)
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
por: Wu, Yanru, et al.
Publicado: (2026)
por: Wu, Yanru, et al.
Publicado: (2026)
Universal Feature Selection for Simultaneous Interpretability of Multitask Datasets
por: Raymond, Matt, et al.
Publicado: (2024)
por: Raymond, Matt, et al.
Publicado: (2024)
Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits
por: Xia, Yu, et al.
Publicado: (2024)
por: Xia, Yu, et al.
Publicado: (2024)
A Continual and Incremental Learning Approach for TinyML On-device Training Using Dataset Distillation and Model Size Adaption
por: Rüb, Marcus, et al.
Publicado: (2024)
por: Rüb, Marcus, et al.
Publicado: (2024)
One Head, Many Models: Cross-Attention Routing for Cost-Aware LLM Selection
por: Pulishetty, Roshini, et al.
Publicado: (2025)
por: Pulishetty, Roshini, et al.
Publicado: (2025)
Sharpness-Aware Parameter Selection for Machine Unlearning
por: Malekmohammadi, Saber, et al.
Publicado: (2025)
por: Malekmohammadi, Saber, et al.
Publicado: (2025)
On the (In)Significance of Feature Selection in High-Dimensional Datasets
por: Neekhra, Bhavesh, et al.
Publicado: (2025)
por: Neekhra, Bhavesh, et al.
Publicado: (2025)
Be Aware of the Neighborhood Effect: Modeling Selection Bias under Interference
por: Li, Haoxuan, et al.
Publicado: (2024)
por: Li, Haoxuan, et al.
Publicado: (2024)
Computation-Aware Kalman Filtering with Model Selection for Neural Dynamics
por: Huml, JR, et al.
Publicado: (2026)
por: Huml, JR, et al.
Publicado: (2026)
TIDES: Implicit Time-Awareness in Selective State Space Models
por: Soydan, Taylan, et al.
Publicado: (2026)
por: Soydan, Taylan, et al.
Publicado: (2026)
CloserMusicDB: A Modern Multipurpose Dataset of High Quality Music
por: Piekarzewicz, Aleksandra, et al.
Publicado: (2024)
por: Piekarzewicz, Aleksandra, et al.
Publicado: (2024)
Ejemplares similares
-
DataMIL: Selecting Data for Robot Imitation Learning with Datamodels
por: Dass, Shivin, et al.
Publicado: (2025) -
Small-to-Large Generalization: Data Influences Models Consistently Across Scale
por: Khaddaj, Alaa, et al.
Publicado: (2025) -
Optimizing ML Training with Metagradient Descent
por: Engstrom, Logan, et al.
Publicado: (2025) -
Data Debiasing with Datamodels (D3M): Improving Subgroup Robustness via Data Selection
por: Jain, Saachi, et al.
Publicado: (2024) -
Attribute-to-Delete: Machine Unlearning via Datamodel Matching
por: Georgiev, Kristian, et al.
Publicado: (2024)