How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Fuente:
arXiv
Guardado en:
| Autores principales: | Bilmes, Jeff A., Bhatt, Gantavya, Das, Arnav M. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Deep Submodular Peripteral Networks
por: Bhatt, Gantavya, et al.
Publicado: (2024)
por: Bhatt, Gantavya, et al.
Publicado: (2024)
Effective Backdoor Mitigation in Vision-Language Models Depends on the Pre-training Objective
por: Verma, Sahil, et al.
Publicado: (2023)
por: Verma, Sahil, et al.
Publicado: (2023)
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
por: Jalali, Mohammad, et al.
Publicado: (2024)
por: Jalali, Mohammad, et al.
Publicado: (2024)
LabelBench: A Comprehensive Framework for Benchmarking Adaptive Label-Efficient Learning
por: Zhang, Jifan, et al.
Publicado: (2023)
por: Zhang, Jifan, et al.
Publicado: (2023)
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)
How Many Images Does It Take? Estimating Imitation Thresholds in Text-to-Image Models
por: Verma, Sahil, et al.
Publicado: (2024)
por: Verma, Sahil, et al.
Publicado: (2024)
Information-Guided Diffusion Sampling for Dataset Distillation
por: Ye, Linfeng, et al.
Publicado: (2025)
por: Ye, Linfeng, et al.
Publicado: (2025)
SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation
por: Zenith, Ayush, et al.
Publicado: (2025)
por: Zenith, Ayush, et al.
Publicado: (2025)
RPN 2: On Interdependence Function Learning Towards Unifying and Advancing CNN, RNN, GNN, and Transformer
por: Zhang, Jiawei
Publicado: (2024)
por: Zhang, Jiawei
Publicado: (2024)
Sharpness-Aware Minimization with Z-Score Gradient Filtering
por: Yun, Vincent-Daniel
Publicado: (2025)
por: Yun, Vincent-Daniel
Publicado: (2025)
How to Evaluate Semantic Communications for Images with ViTScore Metric?
por: Zhu, Tingting, et al.
Publicado: (2023)
por: Zhu, Tingting, et al.
Publicado: (2023)
Towards Formalizing Spuriousness of Biased Datasets Using Partial Information Decomposition
por: Halder, Barproda, et al.
Publicado: (2024)
por: Halder, Barproda, et al.
Publicado: (2024)
Generalized Nested Latent Variable Models for Lossy Coding applied to Wind Turbine Scenarios
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2024)
por: Pérez-Gonzalo, Raül, et al.
Publicado: (2024)
I-Con: A Unifying Framework for Representation Learning
por: Alshammari, Shaden, et al.
Publicado: (2025)
por: Alshammari, Shaden, et al.
Publicado: (2025)
Noise Scheduling as Information-Guided Allocation in Diffusion Training
por: Raya, Gabriel, et al.
Publicado: (2026)
por: Raya, Gabriel, et al.
Publicado: (2026)
Forte : Finding Outliers with Representation Typicality Estimation
por: Ganguly, Debargha, et al.
Publicado: (2024)
por: Ganguly, Debargha, et al.
Publicado: (2024)
Understanding the Role of Equivariance in Self-supervised Learning
por: Wang, Yifei, et al.
Publicado: (2024)
por: Wang, Yifei, et al.
Publicado: (2024)
LVM4CSI: Enabling Direct Application of Pre-Trained Large Vision Models for Wireless Channel Tasks
por: Guo, Jiajia, et al.
Publicado: (2025)
por: Guo, Jiajia, et al.
Publicado: (2025)
Normalized Conditional Mutual Information Surrogate Loss for Deep Neural Classifiers
por: Ye, Linfeng, et al.
Publicado: (2026)
por: Ye, Linfeng, et al.
Publicado: (2026)
Insights from Gradient Dynamics: Gradient Autoscaled Normalization
por: Yun, Vincent-Daniel
Publicado: (2025)
por: Yun, Vincent-Daniel
Publicado: (2025)
Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
por: Xu, Zhaoqi, et al.
Publicado: (2025)
por: Xu, Zhaoqi, et al.
Publicado: (2025)
Poly-View Contrastive Learning
por: Shidani, Amitis, et al.
Publicado: (2024)
por: Shidani, Amitis, et al.
Publicado: (2024)
RPN: Reconciled Polynomial Network Towards Unifying PGMs, Kernel SVMs, MLP and KAN
por: Zhang, Jiawei
Publicado: (2024)
por: Zhang, Jiawei
Publicado: (2024)
Vendi Novelty Scores for Out-of-Distribution Detection
por: Pasarkar, Amey P., et al.
Publicado: (2026)
por: Pasarkar, Amey P., et al.
Publicado: (2026)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
por: Nezhurina, Marianna, et al.
Publicado: (2025)
por: Nezhurina, Marianna, et al.
Publicado: (2025)
The Rate-Distortion-Perception-Classification Tradeoff: Joint Source Coding and Modulation via Inverse-Domain GANs
por: Fang, Junli, et al.
Publicado: (2023)
por: Fang, Junli, et al.
Publicado: (2023)
ASI: Accuracy-Stability Index for Evaluating Deep Learning Models
por: Dai, Wei, et al.
Publicado: (2023)
por: Dai, Wei, et al.
Publicado: (2023)
A Noise is Worth Diffusion Guidance
por: Ahn, Donghoon, et al.
Publicado: (2024)
por: Ahn, Donghoon, et al.
Publicado: (2024)
A Picture is Worth a Thousand Words? An Empirical Study of Aggregation Strategies for Visual Financial Document Retrieval
por: Lim, Ho Hung, et al.
Publicado: (2026)
por: Lim, Ho Hung, et al.
Publicado: (2026)
Universal Representations for Classification-enhanced Lossy Compression
por: Nguyen, Nam
Publicado: (2025)
por: Nguyen, Nam
Publicado: (2025)
Codebook-enabled Generative End-to-end Semantic Communication Powered by Transformer
por: Ye, Peigen, et al.
Publicado: (2024)
por: Ye, Peigen, et al.
Publicado: (2024)
Physics-informed Generalizable Wireless Channel Modeling with Segmentation and Deep Learning: Fundamentals, Methodologies, and Challenges
por: Zhu, Ethan, et al.
Publicado: (2024)
por: Zhu, Ethan, et al.
Publicado: (2024)
Large AI Model-Enabled Generative Semantic Communications for Image Transmission
por: Ma, Qiyu, et al.
Publicado: (2025)
por: Ma, Qiyu, et al.
Publicado: (2025)
Entropy Loss: An Interpretability Amplifier of 3D Object Detection Network for Intelligent Driving
por: Yang, Haobo, et al.
Publicado: (2024)
por: Yang, Haobo, et al.
Publicado: (2024)
ICDM: Interference Cancellation Diffusion Models for Wireless Semantic Communications
por: Wu, Tong, et al.
Publicado: (2025)
por: Wu, Tong, et al.
Publicado: (2025)
Exploiting Information Redundancy in Attention Maps for Extreme Quantization of Vision Transformers
por: Maisonnave, Lucas, et al.
Publicado: (2025)
por: Maisonnave, Lucas, et al.
Publicado: (2025)
Partial Information Decomposition via Normalizing Flows in Latent Gaussian Distributions
por: Zhao, Wenyuan, et al.
Publicado: (2025)
por: Zhao, Wenyuan, et al.
Publicado: (2025)
Deep Generative Sampling in the Dual Divergence Space: A Data-efficient & Interpretative Approach for Generative AI
por: Garg, Sahil, et al.
Publicado: (2024)
por: Garg, Sahil, et al.
Publicado: (2024)
Topology-Aware Exploration of Energy-Based Models Equilibrium: Toric QC-LDPC Codes and Hyperbolic MET QC-LDPC Codes
por: Usatyuk, Vasiliy, et al.
Publicado: (2024)
por: Usatyuk, Vasiliy, et al.
Publicado: (2024)
Visual Language Model based Cross-modal Semantic Communication Systems
por: Jiang, Feibo, et al.
Publicado: (2024)
por: Jiang, Feibo, et al.
Publicado: (2024)
Ejemplares similares
-
Deep Submodular Peripteral Networks
por: Bhatt, Gantavya, et al.
Publicado: (2024) -
Effective Backdoor Mitigation in Vision-Language Models Depends on the Pre-training Objective
por: Verma, Sahil, et al.
Publicado: (2023) -
Conditional Vendi Score: An Information-Theoretic Approach to Diversity Evaluation of Prompt-based Generative Models
por: Jalali, Mohammad, et al.
Publicado: (2024) -
LabelBench: A Comprehensive Framework for Benchmarking Adaptive Label-Efficient Learning
por: Zhang, Jifan, et al.
Publicado: (2023) -
Improving Geo-diversity of Generated Images with Contextualized Vendi Score Guidance
por: Hemmat, Reyhane Askari, et al.
Publicado: (2024)