Effective pruning of web-scale datasets based on complexity of concept clusters
Fuente:
arXiv
Saved in:
| Main Authors: | Abbas, Amro, Rusak, Evgenia, Tirumala, Kushal, Brendel, Wieland, Chaudhuri, Kamalika, Morcos, Ari S. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Low-Pass Filtering Improves Behavioral Alignment of Vision Models
by: Wolff, Max, et al.
Published: (2026)
by: Wolff, Max, et al.
Published: (2026)
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
by: Mayilvahanan, Prasanna, et al.
Published: (2023)
InfoNCE: Identifying the Gap Between Theory and Practice
by: Rusak, Evgenia, et al.
Published: (2024)
by: Rusak, Evgenia, et al.
Published: (2024)
In Search of Forgotten Domain Generalization
by: Mayilvahanan, Prasanna, et al.
Published: (2024)
by: Mayilvahanan, Prasanna, et al.
Published: (2024)
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
by: Mahmoud, Anas, et al.
Published: (2023)
by: Mahmoud, Anas, et al.
Published: (2023)
Scale Alone Does not Improve Mechanistic Interpretability in Vision Models
by: Zimmermann, Roland S., et al.
Published: (2023)
by: Zimmermann, Roland S., et al.
Published: (2023)
Data Redaction from Conditional Generative Models
by: Kong, Zhifeng, et al.
Published: (2023)
by: Kong, Zhifeng, et al.
Published: (2023)
Déjà Vu Memorization in Vision-Language Models
by: Jayaraman, Bargav, et al.
Published: (2024)
by: Jayaraman, Bargav, et al.
Published: (2024)
When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
by: Ramanujan, Vivek, et al.
Published: (2024)
by: Ramanujan, Vivek, et al.
Published: (2024)
CAT: Content-Adaptive Image Tokenization
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models
by: Li, Fanfei, et al.
Published: (2025)
by: Li, Fanfei, et al.
Published: (2025)
Generation is Required for Data-Efficient Perception
by: Brady, Jack, et al.
Published: (2025)
by: Brady, Jack, et al.
Published: (2025)
Measuring Déjà vu Memorization Efficiently
by: Kokhlikyan, Narine, et al.
Published: (2025)
by: Kokhlikyan, Narine, et al.
Published: (2025)
Learning effective pruning at initialization from iterative pruning
by: Liu, Shengkai, et al.
Published: (2024)
by: Liu, Shengkai, et al.
Published: (2024)
Interaction Asymmetry: A General Principle for Learning Composable Abstractions
by: Brady, Jack, et al.
Published: (2024)
by: Brady, Jack, et al.
Published: (2024)
Strawberry detection and counting based on YOLOv7 pruning and information based tracking algorithm
by: Liu, Shiyu, et al.
Published: (2024)
by: Liu, Shiyu, et al.
Published: (2024)
MVTOP: Multi-View Transformer-based Object Pose-Estimation
by: Ranftl, Lukas, et al.
Published: (2025)
by: Ranftl, Lukas, et al.
Published: (2025)
A comparison between humans and AI at recognizing objects in unusual poses
by: Ollikka, Netta, et al.
Published: (2024)
by: Ollikka, Netta, et al.
Published: (2024)
Differentially Private Representation Learning via Image Captioning
by: Sander, Tom, et al.
Published: (2024)
by: Sander, Tom, et al.
Published: (2024)
Neural network relief: a pruning algorithm based on neural activity
by: Dekhovich, Aleksandr, et al.
Published: (2021)
by: Dekhovich, Aleksandr, et al.
Published: (2021)
An accurate detection is not all you need to combat label noise in web-noisy datasets
by: Albert, Paul, et al.
Published: (2024)
by: Albert, Paul, et al.
Published: (2024)
Don't trust your eyes: on the (un)reliability of feature visualizations
by: Geirhos, Robert, et al.
Published: (2023)
by: Geirhos, Robert, et al.
Published: (2023)
Positioning radiata pine branches requiring pruning by drone stereo vision
by: Lin, Yida, et al.
Published: (2026)
by: Lin, Yida, et al.
Published: (2026)
Supporting Vision-Language Model Inference with Confounder-pruning Knowledge Prompt
by: Li, Jiangmeng, et al.
Published: (2022)
by: Li, Jiangmeng, et al.
Published: (2022)
LeCoT: revisiting network architecture for two-view correspondence pruning
by: Dai, Luanyuan, et al.
Published: (2025)
by: Dai, Luanyuan, et al.
Published: (2025)
Interpret, prune and distill Donut : towards lightweight VLMs for VQA on document
by: Mansour, Adnan Ben, et al.
Published: (2025)
by: Mansour, Adnan Ben, et al.
Published: (2025)
An evaluation of Deep Learning based stereo dense matching dataset shift from aerial images and a large scale stereo dataset
by: Wu, Teng, et al.
Published: (2024)
by: Wu, Teng, et al.
Published: (2024)
X-maps: Direct Depth Lookup for Event-based Structured Light Systems
by: Morgenstern, Wieland, et al.
Published: (2024)
by: Morgenstern, Wieland, et al.
Published: (2024)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
by: Zhou, Chunting, et al.
Published: (2024)
by: Zhou, Chunting, et al.
Published: (2024)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
A large-scale dataset for end-to-end table recognition in the wild
by: Yang, Fan, et al.
Published: (2023)
by: Yang, Fan, et al.
Published: (2023)
An aerial color image anomaly dataset for search missions in complex forested terrain
by: Nathan, Rakesh John Amala Arokia, et al.
Published: (2025)
by: Nathan, Rakesh John Amala Arokia, et al.
Published: (2025)
Fit Pixels, Get Labels: Meta-learned Implicit Networks for Image Segmentation
by: Vyas, Kushal, et al.
Published: (2025)
by: Vyas, Kushal, et al.
Published: (2025)
Compact 3D Scene Representation via Self-Organizing Gaussian Grids
by: Morgenstern, Wieland, et al.
Published: (2023)
by: Morgenstern, Wieland, et al.
Published: (2023)
Animating NeRFs from Texture Space: A Framework for Pose-Dependent Rendering of Human Performances
by: Knoll, Paul, et al.
Published: (2023)
by: Knoll, Paul, et al.
Published: (2023)
Brevity is the soul of wit: Pruning long files for code generation
by: Singh, Aaditya K., et al.
Published: (2024)
by: Singh, Aaditya K., et al.
Published: (2024)
Simba: Mamba augmented U-ShiftGCN for Skeletal Action Recognition in Videos
by: Chaudhuri, Soumyabrata, et al.
Published: (2024)
by: Chaudhuri, Soumyabrata, et al.
Published: (2024)
Advanced deep architecture pruning using single filter performance
by: Tzach, Yarden, et al.
Published: (2025)
by: Tzach, Yarden, et al.
Published: (2025)
Deep clustering using adversarial net based clustering loss
by: Lim, Kart-Leong
Published: (2024)
by: Lim, Kart-Leong
Published: (2024)
KOLOMVERSE: Korea open large-scale image dataset for object detection in the maritime universe
by: Nanda, Abhilasha, et al.
Published: (2022)
by: Nanda, Abhilasha, et al.
Published: (2022)
Similar Items
-
Low-Pass Filtering Improves Behavioral Alignment of Vision Models
by: Wolff, Max, et al.
Published: (2026) -
Does CLIP's Generalization Performance Mainly Stem from High Train-Test Similarity?
by: Mayilvahanan, Prasanna, et al.
Published: (2023) -
InfoNCE: Identifying the Gap Between Theory and Practice
by: Rusak, Evgenia, et al.
Published: (2024) -
In Search of Forgotten Domain Generalization
by: Mayilvahanan, Prasanna, et al.
Published: (2024) -
Sieve: Multimodal Dataset Pruning Using Image Captioning Models
by: Mahmoud, Anas, et al.
Published: (2023)