Adapting Self-Supervised Representations as a Latent Space for Efficient Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gui, Ming, Schusterbauer, Johannes, Phan, Timy, Krause, Felix, Susskind, Josh, Bautista, Miguel Angel, Ommer, Björn |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probabilistic Precipitation Nowcasting with Rectified Flow Transformers
by: Schusterbauer, Johannes, et al.
Published: (2026)
by: Schusterbauer, Johannes, et al.
Published: (2026)
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025)
by: Krause, Felix, et al.
Published: (2025)
Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
by: Schusterbauer, Johannes, et al.
Published: (2026)
by: Schusterbauer, Johannes, et al.
Published: (2026)
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
by: Schusterbauer, Johannes, et al.
Published: (2025)
by: Schusterbauer, Johannes, et al.
Published: (2025)
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
by: Stracke, Nick, et al.
Published: (2026)
by: Stracke, Nick, et al.
Published: (2026)
What If : Understanding Motion Through Sparse Interactions
by: Baumann, Stefan Andreas, et al.
Published: (2025)
by: Baumann, Stefan Andreas, et al.
Published: (2025)
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
by: Ma, Pingchuan, et al.
Published: (2025)
by: Ma, Pingchuan, et al.
Published: (2025)
Diffusion Models and Representation Learning: A Survey
by: Fuest, Michael, et al.
Published: (2024)
by: Fuest, Michael, et al.
Published: (2024)
Guiding Token-Sparse Diffusion Models
by: Krause, Felix, et al.
Published: (2026)
by: Krause, Felix, et al.
Published: (2026)
Boosting Latent Diffusion with Flow Matching
by: Schusterbauer, Johannes, et al.
Published: (2023)
by: Schusterbauer, Johannes, et al.
Published: (2023)
CTRLorALTer: Conditional LoRAdapter for Efficient 0-Shot Control & Altering of T2I Models
by: Stracke, Nick, et al.
Published: (2024)
by: Stracke, Nick, et al.
Published: (2024)
Distillation of Diffusion Features for Semantic Correspondence
by: Fundel, Frank, et al.
Published: (2024)
by: Fundel, Frank, et al.
Published: (2024)
DisMo: Disentangled Motion Representations for Open-World Motion Transfer
by: Ressler-Antal, Thomas, et al.
Published: (2025)
by: Ressler-Antal, Thomas, et al.
Published: (2025)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
DepthFM: Fast Monocular Depth Estimation with Flow Matching
by: Gui, Ming, et al.
Published: (2024)
by: Gui, Ming, et al.
Published: (2024)
Unsupervised View-Invariant Human Posture Representation
by: Sardari, Faegheh, et al.
Published: (2021)
by: Sardari, Faegheh, et al.
Published: (2021)
MaskFlow: Discrete Flows For Flexible and Efficient Long Video Generation
by: Fuest, Michael, et al.
Published: (2025)
by: Fuest, Michael, et al.
Published: (2025)
3D Shape Tokenization via Latent Flow Matching
by: Chang, Jen-Hao Rick, et al.
Published: (2024)
by: Chang, Jen-Hao Rick, et al.
Published: (2024)
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP
by: Huang, Chen, et al.
Published: (2024)
by: Huang, Chen, et al.
Published: (2024)
RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
by: Prestel, Ulrich, et al.
Published: (2026)
by: Prestel, Ulrich, et al.
Published: (2026)
Self Supervised Networks for Learning Latent Space Representations of Human Body Scans and Motions
by: Hartman, Emmanuel, et al.
Published: (2024)
by: Hartman, Emmanuel, et al.
Published: (2024)
CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
by: Davtyan, Aram, et al.
Published: (2024)
by: Davtyan, Aram, et al.
Published: (2024)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
[MASK] is All You Need
by: Hu, Vincent Tao, et al.
Published: (2024)
by: Hu, Vincent Tao, et al.
Published: (2024)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
Pseudo-Generalized Dynamic View Synthesis from a Video
by: Zhao, Xiaoming, et al.
Published: (2023)
by: Zhao, Xiaoming, et al.
Published: (2023)
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
by: Yang, Huan, et al.
Published: (2024)
by: Yang, Huan, et al.
Published: (2024)
Normalizing Flows are Capable Generative Models
by: Zhai, Shuangfei, et al.
Published: (2024)
by: Zhai, Shuangfei, et al.
Published: (2024)
Text-Conditional JEPA for Learning Semantically Rich Visual Representations
by: Huang, Chen, et al.
Published: (2026)
by: Huang, Chen, et al.
Published: (2026)
Stable-Pose: Leveraging Transformers for Pose-Guided Text-to-Image Generation
by: Wang, Jiajun, et al.
Published: (2024)
by: Wang, Jiajun, et al.
Published: (2024)
Overcoming the Pitfalls of Vision-Language Model Finetuning for OOD Generalization
by: Zang, Yuhang, et al.
Published: (2024)
by: Zang, Yuhang, et al.
Published: (2024)
Self-Supervised Learning with a Multi-Task Latent Space Objective
by: De Plaen, Pierre-François, et al.
Published: (2026)
by: De Plaen, Pierre-François, et al.
Published: (2026)
Adapting Self-Supervised Learning for Computational Pathology
by: Zimmermann, Eric, et al.
Published: (2024)
by: Zimmermann, Eric, et al.
Published: (2024)
Enhancing Representations through Heterogeneous Self-Supervised Learning
by: Li, Zhong-Yu, et al.
Published: (2023)
by: Li, Zhong-Yu, et al.
Published: (2023)
Continuous, Subject-Specific Attribute Control in T2I Models by Identifying Semantic Directions
by: Baumann, Stefan Andreas, et al.
Published: (2024)
by: Baumann, Stefan Andreas, et al.
Published: (2024)
Latent Drifting in Diffusion Models for Counterfactual Medical Image Synthesis
by: Yeganeh, Yousef, et al.
Published: (2024)
by: Yeganeh, Yousef, et al.
Published: (2024)
World-consistent Video Diffusion with Explicit 3D Modeling
by: Zhang, Qihang, et al.
Published: (2024)
by: Zhang, Qihang, et al.
Published: (2024)
MatSSL: Robust Self-Supervised Representation Learning for Metallographic Image Segmentation
by: Nguyen, Hoang Hai Nam, et al.
Published: (2025)
by: Nguyen, Hoang Hai Nam, et al.
Published: (2025)
Towards Latent Masked Image Modeling for Self-Supervised Visual Representation Learning
by: Wei, Yibing, et al.
Published: (2024)
by: Wei, Yibing, et al.
Published: (2024)
A Generative Framework for Self-Supervised Facial Representation Learning
by: He, Ruian, et al.
Published: (2023)
by: He, Ruian, et al.
Published: (2023)
Similar Items
-
Probabilistic Precipitation Nowcasting with Rectified Flow Transformers
by: Schusterbauer, Johannes, et al.
Published: (2026) -
TREAD: Token Routing for Efficient Architecture-agnostic Diffusion Training
by: Krause, Felix, et al.
Published: (2025) -
Denoising, Fast and Slow: Difficulty-Aware Adaptive Sampling for Image Generation
by: Schusterbauer, Johannes, et al.
Published: (2026) -
Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment
by: Schusterbauer, Johannes, et al.
Published: (2025) -
Learning Long-term Motion Embeddings for Efficient Kinematics Generation
by: Stracke, Nick, et al.
Published: (2026)