[MASK] is All You Need
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Vincent Tao, Ommer, Björn |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Models and Representation Learning: A Survey
by: Fuest, Michael, et al.
Published: (2024)
by: Fuest, Michael, et al.
Published: (2024)
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025)
by: Krause, Felix, et al.
Published: (2025)
Scaling Image Tokenizers with Grouped Spherical Quantization
by: Wang, Jiangtao, et al.
Published: (2024)
by: Wang, Jiangtao, et al.
Published: (2024)
Zoom and Shift are All You Need
by: Qin, Jiahao
Published: (2024)
by: Qin, Jiahao
Published: (2024)
Ideal Registration? Segmentation is All You Need
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Memory augment is All You Need for image restoration
by: Zhang, Xiao Feng, et al.
Published: (2023)
by: Zhang, Xiao Feng, et al.
Published: (2023)
FineVision: Open Data Is All You Need
by: Wiedmann, Luis, et al.
Published: (2025)
by: Wiedmann, Luis, et al.
Published: (2025)
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
by: Baumann, Stefan Andreas, et al.
Published: (2024)
by: Baumann, Stefan Andreas, et al.
Published: (2024)
All You Need in Knowledge Distillation Is a Tailored Coordinate System
by: Zhou, Junjie, et al.
Published: (2024)
by: Zhou, Junjie, et al.
Published: (2024)
Rethinking Deep Clustering Paradigms: Self-Supervision Is All You Need
by: Shaheena, Amal, et al.
Published: (2025)
by: Shaheena, Amal, et al.
Published: (2025)
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation
by: Fuest, Michael, et al.
Published: (2025)
by: Fuest, Michael, et al.
Published: (2025)
Self-supervised Dataset Distillation: A Good Compression Is All You Need
by: Zhou, Muxin, et al.
Published: (2024)
by: Zhou, Muxin, et al.
Published: (2024)
Taxes Are All You Need: Integration of Taxonomical Hierarchy Relationships into the Contrastive Loss
by: Kokilepersaud, Kiran, et al.
Published: (2024)
by: Kokilepersaud, Kiran, et al.
Published: (2024)
CLIP is All You Need for Human-like Semantic Representations in Stable Diffusion
by: Braunstein, Cameron, et al.
Published: (2025)
by: Braunstein, Cameron, et al.
Published: (2025)
Boosting Domain Incremental Learning: Selecting the Optimal Parameters is All You Need
by: Wang, Qiang, et al.
Published: (2025)
by: Wang, Qiang, et al.
Published: (2025)
Anatomy Might Be All You Need: Forecasting What to Do During Surgery
by: Sarwin, Gary, et al.
Published: (2025)
by: Sarwin, Gary, et al.
Published: (2025)
Oasis: One Image is All You Need for Multimodal Instruction Data Synthesis
by: Zhang, Letian, et al.
Published: (2025)
by: Zhang, Letian, et al.
Published: (2025)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
by: Qin, You, et al.
Published: (2024)
by: Qin, You, et al.
Published: (2024)
Fast Wrong-way Cycling Detection in CCTV Videos: Sparse Sampling is All You Need
by: Xu, Jing, et al.
Published: (2024)
by: Xu, Jing, et al.
Published: (2024)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
Off-The-Shelf Image-to-Image Models Are All You Need To Defeat Image Protection Schemes
by: Pleimling, Xavier, et al.
Published: (2026)
by: Pleimling, Xavier, et al.
Published: (2026)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
by: Prestel, Ulrich, et al.
Published: (2026)
by: Prestel, Ulrich, et al.
Published: (2026)
Is Hyperbolic Space All You Need for Medical Anomaly Detection?
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
by: Gonzalez-Jimenez, Alvaro, et al.
Published: (2025)
Text is All You Need for Vision-Language Model Jailbreaking
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
Two Steps Are All You Need: Efficient 3D Point Cloud Anomaly Detection with Consistency Models
by: A, Pranav, et al.
Published: (2026)
by: A, Pranav, et al.
Published: (2026)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
by: Cao, Pu, et al.
Published: (2023)
by: Cao, Pu, et al.
Published: (2023)
Camouflaged Image Synthesis Is All You Need to Boost Camouflaged Detection
by: Zhang, Haichao, et al.
Published: (2023)
by: Zhang, Haichao, et al.
Published: (2023)
Distillation of Diffusion Features for Semantic Correspondence
by: Fundel, Frank, et al.
Published: (2024)
by: Fundel, Frank, et al.
Published: (2024)
Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis
by: Lee, Haeil, et al.
Published: (2024)
by: Lee, Haeil, et al.
Published: (2024)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
Annolid: Annotate, Segment, and Track Anything You Need
by: Yang, Chen, et al.
Published: (2024)
by: Yang, Chen, et al.
Published: (2024)
Compositional Generative Modeling: A Single Model is Not All You Need
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Envisioning the Future, One Step at a Time
by: Baumann, Stefan Andreas, et al.
Published: (2026)
by: Baumann, Stefan Andreas, et al.
Published: (2026)
Procrastination Is All You Need: Exponent Indexed Accumulators for Floating Point, Posits and Logarithmic Numbers
by: Liguori, Vincenzo
Published: (2024)
by: Liguori, Vincenzo
Published: (2024)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
by: Matişan, Răzvan-Andrei, et al.
Published: (2025)
Motivation is Something You Need
by: Acheli, Mehdi, et al.
Published: (2026)
by: Acheli, Mehdi, et al.
Published: (2026)
Does VLM Classification Benefit from LLM Description Semantics?
by: Ma, Pingchuan, et al.
Published: (2024)
by: Ma, Pingchuan, et al.
Published: (2024)
Similar Items
-
Diffusion Models and Representation Learning: A Survey
by: Fuest, Michael, et al.
Published: (2024) -
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025) -
Scaling Image Tokenizers with Grouped Spherical Quantization
by: Wang, Jiangtao, et al.
Published: (2024) -
Zoom and Shift are All You Need
by: Qin, Jiahao
Published: (2024) -
Ideal Registration? Segmentation is All You Need
by: Chen, Xiang, et al.
Published: (2025)