Measuring Style Similarity in Diffusion Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Somepalli, Gowthami, Gupta, Anubhav, Gupta, Kamal, Palta, Shramay, Goldblum, Micah, Geiping, Jonas, Shrivastava, Abhinav, Goldstein, Tom |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
di: Shirkavand, Reza, et al.
Pubblicazione: (2024)
di: Shirkavand, Reza, et al.
Pubblicazione: (2024)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
di: Hayes, Kevin David, et al.
Pubblicazione: (2025)
di: Hayes, Kevin David, et al.
Pubblicazione: (2025)
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
di: Souri, Hossein, et al.
Pubblicazione: (2024)
di: Souri, Hossein, et al.
Pubblicazione: (2024)
CinePile: A Long Video Question Answering Dataset and Benchmark
di: Rawal, Ruchit, et al.
Pubblicazione: (2024)
di: Rawal, Ruchit, et al.
Pubblicazione: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
What do we learn from inverting CLIP models?
di: Kazemi, Hamid, et al.
Pubblicazione: (2024)
di: Kazemi, Hamid, et al.
Pubblicazione: (2024)
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
LEIA: Latent View-invariant Embeddings for Implicit 3D Articulation
di: Swaminathan, Archana, et al.
Pubblicazione: (2024)
di: Swaminathan, Archana, et al.
Pubblicazione: (2024)
The Lie Derivative for Measuring Learned Equivariance
di: Gruver, Nate, et al.
Pubblicazione: (2022)
di: Gruver, Nate, et al.
Pubblicazione: (2022)
Continuous Video Process: Modeling Videos as Continuous Multi-Dimensional Processes for Video Prediction
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
EAGLES: Efficient Accelerated 3D Gaussians with Lightweight EncodingS
di: Girish, Sharath, et al.
Pubblicazione: (2023)
di: Girish, Sharath, et al.
Pubblicazione: (2023)
Adaptive Retention & Correction: Test-Time Training for Continual Learning
di: Chen, Haoran, et al.
Pubblicazione: (2024)
di: Chen, Haoran, et al.
Pubblicazione: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
Gems: Group Emotion Profiling Through Multimodal Situational Understanding
di: Kataria, Anubhav, et al.
Pubblicazione: (2025)
di: Kataria, Anubhav, et al.
Pubblicazione: (2025)
LMD3: Language Model Data Density Dependence
di: Kirchenbauer, John, et al.
Pubblicazione: (2024)
di: Kirchenbauer, John, et al.
Pubblicazione: (2024)
LiFT: A Surprisingly Simple Lightweight Feature Transform for Dense ViT Descriptors
di: Suri, Saksham, et al.
Pubblicazione: (2024)
di: Suri, Saksham, et al.
Pubblicazione: (2024)
Utilization of Neighbor Information for Image Classification with Different Levels of Supervision
di: Jayatilaka, Gihan, et al.
Pubblicazione: (2025)
di: Jayatilaka, Gihan, et al.
Pubblicazione: (2025)
Beyond Visual Similarity: Rule-Guided Multimodal Clustering with explicit domain rules
di: Gupta, Kishor Datta, et al.
Pubblicazione: (2025)
di: Gupta, Kishor Datta, et al.
Pubblicazione: (2025)
Latent-INR: A Flexible Framework for Implicit Representations of Videos with Discriminative Semantics
di: Maiya, Shishira R, et al.
Pubblicazione: (2024)
di: Maiya, Shishira R, et al.
Pubblicazione: (2024)
Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
di: Agarwal, Vatsal, et al.
Pubblicazione: (2025)
di: Agarwal, Vatsal, et al.
Pubblicazione: (2025)
HandsOnVLM: Vision-Language Models for Hand-Object Interaction Prediction
di: Bao, Chen, et al.
Pubblicazione: (2024)
di: Bao, Chen, et al.
Pubblicazione: (2024)
Good Representation, Better Explanation: Role of Convolutional Neural Networks in Transformer-Based Remote Sensing Image Captioning
di: Das, Swadhin, et al.
Pubblicazione: (2025)
di: Das, Swadhin, et al.
Pubblicazione: (2025)
AVID: Adapting Video Diffusion Models to World Models
di: Rigter, Marc, et al.
Pubblicazione: (2024)
di: Rigter, Marc, et al.
Pubblicazione: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
TREND: Tri-teaching for Robust Preference-based Reinforcement Learning with Demonstrations
di: Huang, Shuaiyi, et al.
Pubblicazione: (2025)
di: Huang, Shuaiyi, et al.
Pubblicazione: (2025)
NeRF-Aug: Data Augmentation for Robotics with Neural Radiance Fields
di: Zhu, Eric, et al.
Pubblicazione: (2024)
di: Zhu, Eric, et al.
Pubblicazione: (2024)
Disentanglement in T-space for Faster and Distributed Training of Diffusion Models with Fewer Latent-states
di: Gupta, Samarth, et al.
Pubblicazione: (2025)
di: Gupta, Samarth, et al.
Pubblicazione: (2025)
SIEDD: Shared-Implicit Encoder with Discrete Decoders
di: Rangarajan, Vikram, et al.
Pubblicazione: (2025)
di: Rangarajan, Vikram, et al.
Pubblicazione: (2025)
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
Efficient Continuous Video Flow Model for Video Prediction
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
di: Shrivastava, Gaurav, et al.
Pubblicazione: (2024)
Enhancing Fingerprint Image Synthesis with GANs, Diffusion Models, and Style Transfer Techniques
di: Tang, W., et al.
Pubblicazione: (2024)
di: Tang, W., et al.
Pubblicazione: (2024)
Cluster-Aware Similarity Diffusion for Instance Retrieval
di: Luo, Jifei, et al.
Pubblicazione: (2024)
di: Luo, Jifei, et al.
Pubblicazione: (2024)
Quantifying Cross-Modality Memorization in Vision-Language Models
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation
di: Mejia, Jared, et al.
Pubblicazione: (2024)
di: Mejia, Jared, et al.
Pubblicazione: (2024)
Utilizing Transfer Learning and pre-trained Models for Effective Forest Fire Detection: A Case Study of Uttarakhand
di: Gupta, Hari Prabhat, et al.
Pubblicazione: (2024)
di: Gupta, Hari Prabhat, et al.
Pubblicazione: (2024)
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
di: Mishra, Suyash, et al.
Pubblicazione: (2026)
di: Mishra, Suyash, et al.
Pubblicazione: (2026)
When Generative Augmentation Hurts: A Benchmark Study of GAN and Diffusion Models for Bias Correction in AI Classification Systems
di: Gupta, Shesh Narayan, et al.
Pubblicazione: (2026)
di: Gupta, Shesh Narayan, et al.
Pubblicazione: (2026)
Resolving Spatio-Temporal Entanglement in Video Prediction via Multi-Modal Attention
di: Gupta, Shreyam, et al.
Pubblicazione: (2025)
di: Gupta, Shreyam, et al.
Pubblicazione: (2025)
SIC: Similarity-Based Interpretable Image Classification with Neural Networks
di: Wolf, Tom Nuno, et al.
Pubblicazione: (2025)
di: Wolf, Tom Nuno, et al.
Pubblicazione: (2025)
SiT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant Transformers
di: Ma, Nanye, et al.
Pubblicazione: (2024)
di: Ma, Nanye, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Efficient Fine-Tuning and Concept Suppression for Pruned Diffusion Models
di: Shirkavand, Reza, et al.
Pubblicazione: (2024) -
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
di: Hayes, Kevin David, et al.
Pubblicazione: (2025) -
Generating Potent Poisons and Backdoors from Scratch with Guided Diffusion
di: Souri, Hossein, et al.
Pubblicazione: (2024) -
CinePile: A Long Video Question Answering Dataset and Benchmark
di: Rawal, Ruchit, et al.
Pubblicazione: (2024) -
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)