Discriminative Class Tokens for Text-to-Image Diffusion Models
Fuente:
arXiv
Saved in:
| Main Authors: | Schwartz, Idan, Snæbjarnarson, Vésteinn, Chefer, Hila, Cotterell, Ryan, Belongie, Serge, Wolf, Lior, Benaim, Sagie |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Assessing Neural Network Robustness via Adversarial Pivotal Tuning
by: Christensen, Peter Ebert, et al.
Published: (2022)
by: Christensen, Peter Ebert, et al.
Published: (2022)
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025)
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
by: Zafar, Oz, et al.
Published: (2024)
by: Zafar, Oz, et al.
Published: (2024)
A Meaningful Perturbation Metric for Evaluating Explainability Methods
by: Cohen, Danielle, et al.
Published: (2025)
by: Cohen, Danielle, et al.
Published: (2025)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
by: Yariv, Guy, et al.
Published: (2024)
by: Yariv, Guy, et al.
Published: (2024)
Generating Intermediate Representations for Compositional Text-To-Image Generation
by: Galun, Ran, et al.
Published: (2024)
by: Galun, Ran, et al.
Published: (2024)
FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation
by: Shaulov, Ariel, et al.
Published: (2025)
by: Shaulov, Ariel, et al.
Published: (2025)
MV-RAG: Retrieval Augmented Multiview Diffusion
by: Dayani, Yosef, et al.
Published: (2025)
by: Dayani, Yosef, et al.
Published: (2025)
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
by: Benishu, Omer, et al.
Published: (2026)
by: Benishu, Omer, et al.
Published: (2026)
RAD: Retrieval-Augmented Monocular Metric Depth Estimation for Underrepresented Classes
by: Baltaxe, Michael, et al.
Published: (2026)
by: Baltaxe, Michael, et al.
Published: (2026)
Colored Noise Diffusion Sampling
by: Davidson, Hadar, et al.
Published: (2026)
by: Davidson, Hadar, et al.
Published: (2026)
Coarse-To-Fine Tensor Trains for Compact Visual Representations
by: Loeschcke, Sebastian, et al.
Published: (2024)
by: Loeschcke, Sebastian, et al.
Published: (2024)
VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video Models
by: Chefer, Hila, et al.
Published: (2025)
by: Chefer, Hila, et al.
Published: (2025)
RewardSDS: Aligning Score Distillation via Reward-Weighted Sampling
by: Chachy, Itay, et al.
Published: (2025)
by: Chachy, Itay, et al.
Published: (2025)
Gumbel Counterfactual Generation From Language Models
by: Ravfogel, Shauli, et al.
Published: (2024)
by: Ravfogel, Shauli, et al.
Published: (2024)
TempoControl: Temporal Attention Guidance for Text-to-Video Models
by: Schiber, Shira, et al.
Published: (2025)
by: Schiber, Shira, et al.
Published: (2025)
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
by: Shaulov, Ariel, et al.
Published: (2026)
by: Shaulov, Ariel, et al.
Published: (2026)
Single Image Iterative Subject-driven Generation and Editing
by: Shpitzer, Yair, et al.
Published: (2025)
by: Shpitzer, Yair, et al.
Published: (2025)
GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens
by: Itkin, Roni, et al.
Published: (2026)
by: Itkin, Roni, et al.
Published: (2026)
Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
by: Fiebelman, Gal, et al.
Published: (2025)
by: Fiebelman, Gal, et al.
Published: (2025)
DGD: Dynamic 3D Gaussians Distillation
by: Labe, Isaac, et al.
Published: (2024)
by: Labe, Isaac, et al.
Published: (2024)
Structurally Disentangled Feature Fields Distillation for 3D Understanding and Editing
by: Levy, Yoel, et al.
Published: (2025)
by: Levy, Yoel, et al.
Published: (2025)
Designing a Conditional Prior Distribution for Flow-Based Generative Models
by: Issachar, Noam, et al.
Published: (2025)
by: Issachar, Noam, et al.
Published: (2025)
Stitch: Training-Free Position Control in Multimodal Diffusion Transformers
by: Bader, Jessica, et al.
Published: (2025)
by: Bader, Jessica, et al.
Published: (2025)
Unlearning-based Neural Interpretations
by: Choi, Ching Lam, et al.
Published: (2024)
by: Choi, Ching Lam, et al.
Published: (2024)
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
by: Sreelatha, Silpa Vadakkeeveetil, et al.
Published: (2026)
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
by: Issachar, Noam, et al.
Published: (2025)
by: Issachar, Noam, et al.
Published: (2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
by: Yariv, Guy, et al.
Published: (2025)
by: Yariv, Guy, et al.
Published: (2025)
Still-Moving: Customized Video Generation without Customized Video Data
by: Chefer, Hila, et al.
Published: (2024)
by: Chefer, Hila, et al.
Published: (2024)
SemanticMoments: Training-Free Motion Similarity via Third Moment Features
by: Huberman, Saar, et al.
Published: (2026)
by: Huberman, Saar, et al.
Published: (2026)
Better Language Models Exhibit Higher Visual Alignment
by: Ruthardt, Jona, et al.
Published: (2024)
by: Ruthardt, Jona, et al.
Published: (2024)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
by: Pach, Mateusz, et al.
Published: (2025)
by: Pach, Mateusz, et al.
Published: (2025)
Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes
by: Krakovsky, Shai, et al.
Published: (2025)
by: Krakovsky, Shai, et al.
Published: (2025)
Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
by: Schouten, Marco, et al.
Published: (2026)
by: Schouten, Marco, et al.
Published: (2026)
Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation
by: Shavin, David, et al.
Published: (2026)
by: Shavin, David, et al.
Published: (2026)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
by: Yu, Hong-Tao, et al.
Published: (2025)
by: Yu, Hong-Tao, et al.
Published: (2025)
Stitched Value Model for Diffusion Alignment
by: Go, Hyojun, et al.
Published: (2026)
by: Go, Hyojun, et al.
Published: (2026)
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos
by: Pach, Mateusz, et al.
Published: (2026)
by: Pach, Mateusz, et al.
Published: (2026)
Training-Free Consistent Text-to-Image Generation
by: Tewel, Yoad, et al.
Published: (2024)
by: Tewel, Yoad, et al.
Published: (2024)
Similar Items
-
Assessing Neural Network Robustness via Adversarial Pivotal Tuning
by: Christensen, Peter Ebert, et al.
Published: (2022) -
Taxonomy-Aware Evaluation of Vision-Language Models
by: Snæbjarnarson, Vésteinn, et al.
Published: (2025) -
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
by: Zafar, Oz, et al.
Published: (2024) -
A Meaningful Perturbation Metric for Evaluating Explainability Methods
by: Cohen, Danielle, et al.
Published: (2025) -
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
by: Yariv, Guy, et al.
Published: (2024)