LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images
Fuente:
arXiv
Salvato in:
| Autori principali: | Flora, James, Thopalli, Kowshik, Kulkarni, Akshay R., Wong, Weng-Keen, Liu, Shusen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
Speeding Up Image Classifiers with Little Companions
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
On the Use of Anchoring for Training Vision Models
di: Narayanaswamy, Vivek, et al.
Pubblicazione: (2024)
di: Narayanaswamy, Vivek, et al.
Pubblicazione: (2024)
DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
di: Subramanyam, Rakshith, et al.
Pubblicazione: (2024)
di: Subramanyam, Rakshith, et al.
Pubblicazione: (2024)
Interpretability-Guided Test-Time Adversarial Defense
di: Kulkarni, Akshay, et al.
Pubblicazione: (2024)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2024)
Leveraging Registers in Vision Transformers for Robust Adaptation
di: Yellapragada, Srikar, et al.
Pubblicazione: (2025)
di: Yellapragada, Srikar, et al.
Pubblicazione: (2025)
Beyond Top Activations: Efficient and Reliable Crowdsourced Evaluation of Automated Interpretability
di: Oikarinen, Tuomas, et al.
Pubblicazione: (2025)
di: Oikarinen, Tuomas, et al.
Pubblicazione: (2025)
Orchid: Image Latent Diffusion for Joint Appearance and Geometry Generation
di: Krishnan, Akshay, et al.
Pubblicazione: (2025)
di: Krishnan, Akshay, et al.
Pubblicazione: (2025)
SOOD-ImageNet: a Large-Scale Dataset for Semantic Out-Of-Distribution Image Classification and Semantic Segmentation
di: Bacchin, Alberto, et al.
Pubblicazione: (2024)
di: Bacchin, Alberto, et al.
Pubblicazione: (2024)
Interpreting Neurons in Deep Vision Networks with Language Models
di: Bai, Nicholas, et al.
Pubblicazione: (2024)
di: Bai, Nicholas, et al.
Pubblicazione: (2024)
Interpretable Generative Models through Post-hoc Concept Bottlenecks
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025)
Visual Exploration of Feature Relationships in Sparse Autoencoders with Curated Concepts
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
di: Yan, Xinyuan, et al.
Pubblicazione: (2025)
Scaling Large Motion Models with Million-Level Human Motions
di: Wang, Ye, et al.
Pubblicazione: (2024)
di: Wang, Ye, et al.
Pubblicazione: (2024)
Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation
di: Zhou, Yufan, et al.
Pubblicazione: (2024)
di: Zhou, Yufan, et al.
Pubblicazione: (2024)
Towards Resource-Efficient Streaming of Large-Scale Medical Image Datasets for Deep Learning
di: Kulkarni, Pranav, et al.
Pubblicazione: (2023)
di: Kulkarni, Pranav, et al.
Pubblicazione: (2023)
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
di: Wang, Yuhan, et al.
Pubblicazione: (2025)
BIKED++: A Multimodal Dataset of 1.4 Million Bicycle Image and Parametric CAD Designs
di: Regenwetter, Lyle, et al.
Pubblicazione: (2024)
di: Regenwetter, Lyle, et al.
Pubblicazione: (2024)
StableSemantics: A Synthetic Language-Vision Dataset of Semantic Representations in Naturalistic Images
di: Zawar, Rushikesh, et al.
Pubblicazione: (2024)
di: Zawar, Rushikesh, et al.
Pubblicazione: (2024)
TIP-I2V: A Million-Scale Real Text and Image Prompt Dataset for Image-to-Video Generation
di: Wang, Wenhao, et al.
Pubblicazione: (2024)
di: Wang, Wenhao, et al.
Pubblicazione: (2024)
EmerDiff: Emerging Pixel-level Semantic Knowledge in Diffusion Models
di: Namekata, Koichi, et al.
Pubblicazione: (2024)
di: Namekata, Koichi, et al.
Pubblicazione: (2024)
ZeroDiff: Solidified Visual-Semantic Correlation in Zero-Shot Learning
di: Ye, Zihan, et al.
Pubblicazione: (2024)
di: Ye, Zihan, et al.
Pubblicazione: (2024)
X-Mark: Saliency-Guided Robust Dataset Ownership Verification for Medical Imaging
di: Kulkarni, Pranav, et al.
Pubblicazione: (2026)
di: Kulkarni, Pranav, et al.
Pubblicazione: (2026)
Semantic Augmentation in Images using Language
di: Yerramilli, Sahiti, et al.
Pubblicazione: (2024)
di: Yerramilli, Sahiti, et al.
Pubblicazione: (2024)
SemDiff: Generating Natural Unrestricted Adversarial Examples via Semantic Attributes Optimization in Diffusion Models
di: Dai, Zeyu, et al.
Pubblicazione: (2025)
di: Dai, Zeyu, et al.
Pubblicazione: (2025)
DiffHarmony: Latent Diffusion Model Meets Image Harmonization
di: Zhou, Pengfei, et al.
Pubblicazione: (2024)
di: Zhou, Pengfei, et al.
Pubblicazione: (2024)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
di: Wallingford, Matthew, et al.
Pubblicazione: (2024)
DisCo-Diff: Enhancing Continuous Diffusion Models with Discrete Latents
di: Xu, Yilun, et al.
Pubblicazione: (2024)
di: Xu, Yilun, et al.
Pubblicazione: (2024)
VideoUFO: A Million-Scale User-Focused Dataset for Text-to-Video Generation
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
di: Wang, Wenhao, et al.
Pubblicazione: (2025)
CubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generation
di: Kalischek, Nikolai, et al.
Pubblicazione: (2025)
di: Kalischek, Nikolai, et al.
Pubblicazione: (2025)
DiffEditor: Boosting Accuracy and Flexibility on Diffusion-based Image Editing
di: Mou, Chong, et al.
Pubblicazione: (2024)
di: Mou, Chong, et al.
Pubblicazione: (2024)
Enhancing Accuracy and Parameter-Efficiency of Neural Representations for Network Parameterization
di: Choi, Hongjun, et al.
Pubblicazione: (2024)
di: Choi, Hongjun, et al.
Pubblicazione: (2024)
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
di: Fang, Rongyao, et al.
Pubblicazione: (2025)
di: Fang, Rongyao, et al.
Pubblicazione: (2025)
FRACTAL: An Ultra-Large-Scale Aerial Lidar Dataset for 3D Semantic Segmentation of Diverse Landscapes
di: Gaydon, Charles, et al.
Pubblicazione: (2024)
di: Gaydon, Charles, et al.
Pubblicazione: (2024)
SLICE: Semantic Latent Injection via Compartmentalized Embedding for Image Watermarking
di: Gao, Zheng, et al.
Pubblicazione: (2026)
di: Gao, Zheng, et al.
Pubblicazione: (2026)
Understanding Bias in Large-Scale Visual Datasets
di: Zeng, Boya, et al.
Pubblicazione: (2024)
di: Zeng, Boya, et al.
Pubblicazione: (2024)
DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models
di: Bertazzini, Giulia, et al.
Pubblicazione: (2025)
di: Bertazzini, Giulia, et al.
Pubblicazione: (2025)
Pre-training of Lightweight Vision Transformers on Small Datasets with Minimally Scaled Images
di: Tan, Jen Hong
Pubblicazione: (2024)
di: Tan, Jen Hong
Pubblicazione: (2024)
Quilt-1M: One Million Image-Text Pairs for Histopathology
di: Ikezogwo, Wisdom Oluchi, et al.
Pubblicazione: (2023)
di: Ikezogwo, Wisdom Oluchi, et al.
Pubblicazione: (2023)
ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
di: Krumpl, Gerhard, et al.
Pubblicazione: (2026)
di: Krumpl, Gerhard, et al.
Pubblicazione: (2026)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
di: Nezhurina, Marianna, et al.
Pubblicazione: (2025)
di: Nezhurina, Marianna, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
di: Kulkarni, Akshay, et al.
Pubblicazione: (2025) -
Speeding Up Image Classifiers with Little Companions
di: Liu, Yang, et al.
Pubblicazione: (2024) -
On the Use of Anchoring for Training Vision Models
di: Narayanaswamy, Vivek, et al.
Pubblicazione: (2024) -
DECIDER: Leveraging Foundation Model Priors for Improved Model Failure Detection and Explanation
di: Subramanyam, Rakshith, et al.
Pubblicazione: (2024) -
Interpretability-Guided Test-Time Adversarial Defense
di: Kulkarni, Akshay, et al.
Pubblicazione: (2024)