USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Chu, Xiangxiang, Li, Renda, Wang, Yong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Understanding Self-Supervised Pretraining with Part-Aware Representation Learning
by: Zhu, Jie, et al.
Published: (2023)
by: Zhu, Jie, et al.
Published: (2023)
SSPFormer: Self-Supervised Pretrained Transformer for MRI Images
by: Li, Jingkai, et al.
Published: (2026)
by: Li, Jingkai, et al.
Published: (2026)
LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling
by: Li, Huaqiu, et al.
Published: (2025)
by: Li, Huaqiu, et al.
Published: (2025)
UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
by: Zhang, Xiao, et al.
Published: (2025)
by: Zhang, Xiao, et al.
Published: (2025)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)
by: Lei, Jiachen, et al.
Published: (2025)
GenView: Enhancing View Quality with Pretrained Generative Model for Self-Supervised Learning
by: Li, Xiaojie, et al.
Published: (2024)
by: Li, Xiaojie, et al.
Published: (2024)
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
by: Wang, Zengbin, et al.
Published: (2026)
by: Wang, Zengbin, et al.
Published: (2026)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
by: Chu, Xiangxiang, et al.
Published: (2024)
by: Chu, Xiangxiang, et al.
Published: (2024)
FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow
by: Li, Hangyu, et al.
Published: (2024)
by: Li, Hangyu, et al.
Published: (2024)
SSP-RACL: Classification of Noisy Fundus Images with Self-Supervised Pretraining and Robust Adaptive Credal Loss
by: Ye, Mengwen, et al.
Published: (2024)
by: Ye, Mengwen, et al.
Published: (2024)
ARVideo: Autoregressive Pretraining for Self-Supervised Video Representation Learning
by: Ren, Sucheng, et al.
Published: (2024)
by: Ren, Sucheng, et al.
Published: (2024)
Self-Supervised Pretraining for Fine-Grained Plankton Recognition
by: Kareinen, Joona, et al.
Published: (2025)
by: Kareinen, Joona, et al.
Published: (2025)
UniTable: Towards a Unified Framework for Table Recognition via Self-Supervised Pretraining
by: Peng, ShengYun, et al.
Published: (2024)
by: Peng, ShengYun, et al.
Published: (2024)
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens
by: Li, Zekun, et al.
Published: (2026)
by: Li, Zekun, et al.
Published: (2026)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
by: Li, Mingxing, et al.
Published: (2025)
by: Li, Mingxing, et al.
Published: (2025)
Spectral-Spatial Self-Supervised Learning for Few-Shot Hyperspectral Image Classification
by: Chen, Wenchen, et al.
Published: (2025)
by: Chen, Wenchen, et al.
Published: (2025)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
by: Wang, Jiepeng, et al.
Published: (2025)
by: Wang, Jiepeng, et al.
Published: (2025)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
by: Song, Xinyang, et al.
Published: (2025)
by: Song, Xinyang, et al.
Published: (2025)
Image Clustering Algorithm Based on Self-Supervised Pretrained Models and Latent Feature Distribution Optimization
by: Zhu, Qiuyu, et al.
Published: (2024)
by: Zhu, Qiuyu, et al.
Published: (2024)
Emerging Properties in Unified Multimodal Pretraining
by: Deng, Chaorui, et al.
Published: (2025)
by: Deng, Chaorui, et al.
Published: (2025)
Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
by: Claessens, Cris, et al.
Published: (2025)
by: Claessens, Cris, et al.
Published: (2025)
Unified Reward Model for Multimodal Understanding and Generation
by: Wang, Yibin, et al.
Published: (2025)
by: Wang, Yibin, et al.
Published: (2025)
Self-Supervised Pretraining for Aerial Road Extraction
by: Polley, Rupert, et al.
Published: (2025)
by: Polley, Rupert, et al.
Published: (2025)
SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation
by: Chen, Zhiquan, et al.
Published: (2026)
by: Chen, Zhiquan, et al.
Published: (2026)
ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining
by: Ma, Qi, et al.
Published: (2024)
by: Ma, Qi, et al.
Published: (2024)
Demographic-Aware Self-Supervised Anomaly Detection Pretraining for Equitable Rare Cardiac Diagnosis
by: Huang, Chaoqin, et al.
Published: (2026)
by: Huang, Chaoqin, et al.
Published: (2026)
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
by: Li, Haoyuan, et al.
Published: (2025)
by: Li, Haoyuan, et al.
Published: (2025)
Text-Guided Channel Perturbation and Pretrained Knowledge Integration for Unified Multi-Modality Image Fusion
by: Li, Xilai, et al.
Published: (2025)
by: Li, Xilai, et al.
Published: (2025)
Using Deep Learning Models Pretrained by Self-Supervised Learning for Protein Localization
by: Isselmann, Ben, et al.
Published: (2026)
by: Isselmann, Ben, et al.
Published: (2026)
UNO: Unified Self-Supervised Monocular Odometry for Platform-Agnostic Deployment
by: Zhao, Wentao, et al.
Published: (2025)
by: Zhao, Wentao, et al.
Published: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
by: Liu, Zeyu, et al.
Published: (2026)
by: Liu, Zeyu, et al.
Published: (2026)
Pretraining Deformable Image Registration Networks with Random Images
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
SketchTriplet: Self-Supervised Scenarized Sketch-Text-Image Triplet Generation
by: Wu, Zhenbei, et al.
Published: (2024)
by: Wu, Zhenbei, et al.
Published: (2024)
Understand Before You Generate: Self-Guided Training for Autoregressive Image Generation
by: Yue, Xiaoyu, et al.
Published: (2025)
by: Yue, Xiaoyu, et al.
Published: (2025)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
by: Hao, Xiangzhao, et al.
Published: (2026)
by: Hao, Xiangzhao, et al.
Published: (2026)
MMGenBench: Fully Automatically Evaluating LMMs from the Text-to-Image Generation Perspective
by: Huang, Hailang, et al.
Published: (2024)
by: Huang, Hailang, et al.
Published: (2024)
Faster Training, Fewer Labels: Self-Supervised Pretraining for Fine-Grained BEV Segmentation
by: Busch, Daniel, et al.
Published: (2026)
by: Busch, Daniel, et al.
Published: (2026)
A Unified Framework for Semi-Supervised Image Segmentation and Registration
by: Li, Ruizhe, et al.
Published: (2025)
by: Li, Ruizhe, et al.
Published: (2025)
Mitigating Overfitting in Medical Imaging: Self-Supervised Pretraining vs. ImageNet Transfer Learning for Dermatological Diagnosis
by: Matas, Iván, et al.
Published: (2025)
by: Matas, Iván, et al.
Published: (2025)
Dual Diffusion for Unified Image Generation and Understanding
by: Li, Zijie, et al.
Published: (2024)
by: Li, Zijie, et al.
Published: (2024)
Similar Items
-
Understanding Self-Supervised Pretraining with Part-Aware Representation Learning
by: Zhu, Jie, et al.
Published: (2023) -
SSPFormer: Self-Supervised Pretrained Transformer for MRI Images
by: Li, Jingkai, et al.
Published: (2026) -
LD-RPS: Zero-Shot Unified Image Restoration via Latent Diffusion Recurrent Posterior Sampling
by: Li, Huaqiu, et al.
Published: (2025) -
UPRE: Zero-Shot Domain Adaptation for Object Detection via Unified Prompt and Representation Enhancement
by: Zhang, Xiao, et al.
Published: (2025) -
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
by: Lei, Jiachen, et al.
Published: (2025)