CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
Fuente:
arXiv
Saved in:
| Main Authors: | Braunstein, Cameron, Toneva, Mariya, Ilg, Eddy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SLayR: Scene Layout Generation with Rectified Flow
by: Braunstein, Cameron, et al.
Published: (2024)
by: Braunstein, Cameron, et al.
Published: (2024)
Quantum-Hybrid Stereo Matching With Nonlinear Regularization and Spatial Pyramids
by: Braunstein, Cameron, et al.
Published: (2023)
by: Braunstein, Cameron, et al.
Published: (2023)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
by: Ahmed, Noor, et al.
Published: (2025)
by: Ahmed, Noor, et al.
Published: (2025)
Zoom and Shift are All You Need
by: Qin, Jiahao
Published: (2024)
by: Qin, Jiahao
Published: (2024)
Ideal Registration? Segmentation is All You Need
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
FineVision: Open Data Is All You Need
by: Wiedmann, Luis, et al.
Published: (2025)
by: Wiedmann, Luis, et al.
Published: (2025)
Memory augment is All You Need for image restoration
by: Zhang, Xiao Feng, et al.
Published: (2023)
by: Zhang, Xiao Feng, et al.
Published: (2023)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
by: Cao, Pu, et al.
Published: (2023)
by: Cao, Pu, et al.
Published: (2023)
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
by: Shaheena, Amal, et al.
Published: (2025)
by: Shaheena, Amal, et al.
Published: (2025)
All You Need in Knowledge Distillation Is a Tailored Coordinate System
by: Zhou, Junjie, et al.
Published: (2024)
by: Zhou, Junjie, et al.
Published: (2024)
Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis
by: Lee, Haeil, et al.
Published: (2024)
by: Lee, Haeil, et al.
Published: (2024)
Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
by: Sarwin, Gary, et al.
Published: (2025)
by: Sarwin, Gary, et al.
Published: (2025)
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
by: Zhang, Letian, et al.
Published: (2025)
by: Zhang, Letian, et al.
Published: (2025)
Self-supervised Dataset Distillation: A Good Compression Is All You Need
by: Zhou, Muxin, et al.
Published: (2024)
by: Zhou, Muxin, et al.
Published: (2024)
Taxes Are All You Need: Integration of Taxonomical Hierarchy Relationships into the Contrastive Loss
by: Kokilepersaud, Kiran, et al.
Published: (2024)
by: Kokilepersaud, Kiran, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
Off-The-Shelf Image-to-Image Models Are All You Need To Defeat Image Protection Schemes
by: Pleimling, Xavier, et al.
Published: (2026)
by: Pleimling, Xavier, et al.
Published: (2026)
Fast Wrong-way Cycling Detection in CCTV Videos: Sparse Sampling is All You Need
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Quantifying and Enabling the Interpretability of CLIP-like Models
by: Madasu, Avinash, et al.
Published: (2024)
by: Madasu, Avinash, et al.
Published: (2024)
Memorization In Stable Diffusion Is Unexpectedly Driven by CLIP Embeddings
by: Kim, Bumjun, et al.
Published: (2026)
by: Kim, Bumjun, et al.
Published: (2026)
Is Hyperbolic Space All You Need for Medical Anomaly Detection?
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
Text is All You Need for Vision-Language Model Jailbreaking
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
by: Zhou, Qiongyi, et al.
Published: (2024)
by: Zhou, Qiongyi, et al.
Published: (2024)
Two Steps Are All You Need: Efficient 3D Point Cloud Anomaly Detection with Consistency Models
by: A, Pranav, et al.
Published: (2026)
by: A, Pranav, et al.
Published: (2026)
Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
by: Zhang, Haichao, et al.
Published: (2023)
by: Zhang, Haichao, et al.
Published: (2023)
CLIP-Guided Unsupervised Semantic-Aware Exposure Correction
by: Wu, Puzhen, et al.
Published: (2026)
by: Wu, Puzhen, et al.
Published: (2026)
Demonstrating and Reducing Shortcuts in Vision-Language Representation Learning
by: Bleeker, Maurits, et al.
Published: (2024)
by: Bleeker, Maurits, et al.
Published: (2024)
Knowledge-Base based Semantic Image Transmission Using CLIP
by: Li, Chongyang, et al.
Published: (2025)
by: Li, Chongyang, et al.
Published: (2025)
The One Where They Brain-Tune for Social Cognition: Multi-Modal Brain-Tuning on Friends
by: Policzer, Nico, et al.
Published: (2025)
by: Policzer, Nico, et al.
Published: (2025)
Hierarchical Representation Matching for CLIP-based Class-Incremental Learning
by: Wen, Zhen-Hao, et al.
Published: (2025)
by: Wen, Zhen-Hao, et al.
Published: (2025)
Leveraging Cross-Modal Neighbor Representation for Improved CLIP Classification
by: Yi, Chao, et al.
Published: (2024)
by: Yi, Chao, et al.
Published: (2024)
Interpreting CLIP's Image Representation via Text-Based Decomposition
by: Gandelsman, Yossi, et al.
Published: (2023)
by: Gandelsman, Yossi, et al.
Published: (2023)
EDT: An Efficient Diffusion Transformer Framework Inspired by Human-like Sketching
by: Chen, Xinwang, et al.
Published: (2024)
by: Chen, Xinwang, et al.
Published: (2024)
Annolid: Annotate, Segment, and Track Anything You Need
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
CLIPin: A Non-contrastive Plug-in to CLIP for Multimodal Semantic Alignment
by: Yang, Shengzhu, et al.
Published: (2025)
by: Yang, Shengzhu, et al.
Published: (2025)
Similar Items
-
SLayR: Scene Layout Generation with Rectified Flow
by: Braunstein, Cameron, et al.
Published: (2024) -
Quantum-Hybrid Stereo Matching With Nonlinear Regularization and Spatial Pyramids
by: Braunstein, Cameron, et al.
Published: (2023) -
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025) -
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024) -
3DFroMLLM: 3D Prototype Generation only from Pretrained Multimodal LLMs
by: Ahmed, Noor, et al.
Published: (2025)