One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Yuan, Chen, Chen, Chen, Tianrong, Gu, Jiatao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Normalizing Flows with Iterative Denoising
by: Chen, Tianrong, et al.
Published: (2026)
by: Chen, Tianrong, et al.
Published: (2026)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
by: Guo, Qingpei, et al.
Published: (2024)
by: Guo, Qingpei, et al.
Published: (2024)
Is One GPU Enough? Pushing Image Generation at Higher-Resolutions with Foundation Models
by: Tragakis, Athanasios, et al.
Published: (2024)
by: Tragakis, Athanasios, et al.
Published: (2024)
IPCV: Information-Preserving Compression for MLLM Visual Encoders
by: Chen, Yuan, et al.
Published: (2025)
by: Chen, Yuan, et al.
Published: (2025)
STARFlow: Scaling Latent Normalizing Flows for High-resolution Image Synthesis
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
Augmented Conditioning Is Enough For Effective Training Image Generation
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering
by: Chen, Yuhan, et al.
Published: (2024)
by: Chen, Yuhan, et al.
Published: (2024)
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
by: Kumar, Ashwin, et al.
Published: (2026)
by: Kumar, Ashwin, et al.
Published: (2026)
One Pool Is Not Enough: Multi-Cluster Memory for Practical Test-Time Adaptation
by: Tseng, Yu-Wen, et al.
Published: (2026)
by: Tseng, Yu-Wen, et al.
Published: (2026)
Layered Diffusion Model for One-Shot High Resolution Text-to-Image Synthesis
by: Khwaja, Emaad, et al.
Published: (2024)
by: Khwaja, Emaad, et al.
Published: (2024)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
TypeScore: A Text Fidelity Metric for Text-to-Image Generative Models
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
by: Sampaio, Georgia Gabriela, et al.
Published: (2024)
Generation of Heterogeneous PET Images from Uniform Organ Activity Maps Using a Pretrained Domain-Adapted Diffusion Model
by: Li, Suya, et al.
Published: (2026)
by: Li, Suya, et al.
Published: (2026)
DELST: Dual Entailment Learning for Hyperbolic Image-Gene Pretraining in Spatial Transcriptomics
by: Chen, Xulin, et al.
Published: (2025)
by: Chen, Xulin, et al.
Published: (2025)
RAP: Retrieve, Adapt, and Prompt-Fit for Training-Free Few-Shot Medical Image Segmentation
by: Mao, Zhihao, et al.
Published: (2026)
by: Mao, Zhihao, et al.
Published: (2026)
3D MRI Image Pretraining via Controllable 2D Slice Navigation Task
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
by: Chen, Jiuhai, et al.
Published: (2024)
by: Chen, Jiuhai, et al.
Published: (2024)
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
by: Rajagopalan, Sudarshan, et al.
Published: (2025)
by: Rajagopalan, Sudarshan, et al.
Published: (2025)
Patch-enhanced Mask Encoder Prompt Image Generation
by: Xu, Shusong, et al.
Published: (2024)
by: Xu, Shusong, et al.
Published: (2024)
LMFusion: Adapting Pretrained Language Models for Multimodal Generation
by: Shi, Weijia, et al.
Published: (2024)
by: Shi, Weijia, et al.
Published: (2024)
General Purpose Image Encoder DINOv2 for Medical Image Registration
by: Song, Xinrui, et al.
Published: (2024)
by: Song, Xinrui, et al.
Published: (2024)
One-Step is Enough: Sparse Autoencoders for Text-to-Image Diffusion Models
by: Surkov, Viacheslav, et al.
Published: (2024)
by: Surkov, Viacheslav, et al.
Published: (2024)
MISS: A Generative Pretraining and Finetuning Approach for Med-VQA
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
by: Jia, Ziheng, et al.
Published: (2025)
by: Jia, Ziheng, et al.
Published: (2025)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
by: Han, Donghoon, et al.
Published: (2026)
by: Han, Donghoon, et al.
Published: (2026)
Representations of Text and Images Align From Layer One
by: Wybitul, Evžen, et al.
Published: (2026)
by: Wybitul, Evžen, et al.
Published: (2026)
SEDEG:Sequential Enhancement of Decoder and Encoder's Generality for Class Incremental Learning with Small Memory
by: Chen, Hongyang, et al.
Published: (2025)
by: Chen, Hongyang, et al.
Published: (2025)
MoVE-KD: Knowledge Distillation for VLMs with Mixture of Visual Encoders
by: Cao, Jiajun, et al.
Published: (2025)
by: Cao, Jiajun, et al.
Published: (2025)
Denoising Score Distillation: From Noisy Diffusion Pretraining to One-Step High-Quality Generation
by: Chen, Tianyu, et al.
Published: (2025)
by: Chen, Tianyu, et al.
Published: (2025)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
by: Liu, Xuewen, et al.
Published: (2026)
by: Liu, Xuewen, et al.
Published: (2026)
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
by: Lin, Bin, et al.
Published: (2025)
by: Lin, Bin, et al.
Published: (2025)
TIER: Text-Image Encoder-based Regression for AIGC Image Quality Assessment
by: Yuan, Jiquan, et al.
Published: (2024)
by: Yuan, Jiquan, et al.
Published: (2024)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
Skrr: Skip and Re-use Text Encoder Layers for Memory Efficient Text-to-Image Generation
by: Seo, Hoigi, et al.
Published: (2025)
by: Seo, Hoigi, et al.
Published: (2025)
WSI-VQA: Interpreting Whole Slide Images by Generative Visual Question Answering
by: Chen, Pingyi, et al.
Published: (2024)
by: Chen, Pingyi, et al.
Published: (2024)
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
by: Cao, Zhuo, et al.
Published: (2025)
by: Cao, Zhuo, et al.
Published: (2025)
Similar Items
-
Normalizing Flows with Iterative Denoising
by: Chen, Tianrong, et al.
Published: (2026) -
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026) -
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
by: Guo, Qingpei, et al.
Published: (2024) -
Is One GPU Enough? Pushing Image Generation at Higher-Resolutions with Foundation Models
by: Tragakis, Athanasios, et al.
Published: (2024) -
IPCV: Information-Preserving Compression for MLLM Visual Encoders
by: Chen, Yuan, et al.
Published: (2025)