Representation Alignment for Just Image Transformers is not Easier than You Think
Fuente:
arXiv
Saved in:
| Main Authors: | Shin, Jaeyo, Kim, Jiwook, Shim, Hyunjung |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
by: Kim, Jiwook, et al.
Published: (2024)
by: Kim, Jiwook, et al.
Published: (2024)
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024)
by: Yu, Sihyun, et al.
Published: (2024)
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2025)
by: Choi, Hojun, et al.
Published: (2025)
Aligning Text to Image in Diffusion Models is Easier Than You Think
by: Lee, Jaa-Yeon, et al.
Published: (2025)
by: Lee, Jaa-Yeon, et al.
Published: (2025)
Rethinking the Use of Vision Transformers for AI-Generated Image Detection
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
by: Sun, Zening, et al.
Published: (2026)
by: Sun, Zening, et al.
Published: (2026)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
Scribble-Guided Diffusion for Training-free Text-to-Image Generation
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Directional Textual Inversion for Personalized Text-to-Image Generation
by: Kim, Kunhee, et al.
Published: (2025)
by: Kim, Kunhee, et al.
Published: (2025)
DGQ: Distribution-Aware Group Quantization for Text-to-Image Diffusion Models
by: Ryu, Hyogon, et al.
Published: (2025)
by: Ryu, Hyogon, et al.
Published: (2025)
MagiCapture: High-Resolution Multi-Concept Portrait Customization
by: Hyung, Junha, et al.
Published: (2023)
by: Hyung, Junha, et al.
Published: (2023)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
by: Garcia, Gonzalo Martin, et al.
Published: (2024)
Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
by: Chen, Liang, et al.
Published: (2025)
by: Chen, Liang, et al.
Published: (2025)
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models
by: Park, NaHyeon, et al.
Published: (2025)
by: Park, NaHyeon, et al.
Published: (2025)
Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
by: Wu, Ge, et al.
Published: (2025)
by: Wu, Ge, et al.
Published: (2025)
I0T: Embedding Standardization Method Towards Zero Modality Gap
by: An, Na Min, et al.
Published: (2024)
by: An, Na Min, et al.
Published: (2024)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
by: Yu, Seungjun, et al.
Published: (2025)
by: Yu, Seungjun, et al.
Published: (2025)
Learning to Think in Physics: Breaking Shortcut Learning in Scientific Diffusion via Representation Alignment
by: Jia, Haozhe, et al.
Published: (2026)
by: Jia, Haozhe, et al.
Published: (2026)
Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps
by: Kim, Jeeyung, et al.
Published: (2024)
by: Kim, Jeeyung, et al.
Published: (2024)
Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
by: Jung, Raehyuk, et al.
Published: (2025)
by: Jung, Raehyuk, et al.
Published: (2025)
EasyVideoR1: Easier RL for Video Understanding
by: Qin, Chuanyu, et al.
Published: (2026)
by: Qin, Chuanyu, et al.
Published: (2026)
Addressing Image Hallucination in Text-to-Image Generation through Factual Image Retrieval
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
Masked Generative Transformer Is What You Need for Image Editing
by: Chow, Wei, et al.
Published: (2026)
by: Chow, Wei, et al.
Published: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
Preserving Pre-trained Representation Space: On Effectiveness of Prefix-tuning for Large Multi-modal Models
by: Kim, Donghoon, et al.
Published: (2024)
by: Kim, Donghoon, et al.
Published: (2024)
Classifier-guided CLIP Distillation for Unsupervised Multi-label Classification
by: Kim, Dongseob, et al.
Published: (2025)
by: Kim, Dongseob, et al.
Published: (2025)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
by: Lee, Seungho, et al.
Published: (2024)
by: Lee, Seungho, et al.
Published: (2024)
DIAR: Deep Image Alignment and Reconstruction using Swin Transformers
by: Kwiatkowski, Monika, et al.
Published: (2023)
by: Kwiatkowski, Monika, et al.
Published: (2023)
Equivariant Latent Alignment via Flow Matching under Group Symmetries
by: Kim, Sunghyun, et al.
Published: (2026)
by: Kim, Sunghyun, et al.
Published: (2026)
Real-Time Long Horizon Air Quality Forecasting via Group-Relative Policy Optimization
by: Kang, Inha, et al.
Published: (2025)
by: Kang, Inha, et al.
Published: (2025)
Reliable Thinking with Images
by: Li, Haobin, et al.
Published: (2026)
by: Li, Haobin, et al.
Published: (2026)
OCT Data is All You Need: How Vision Transformers with and without Pre-training Benefit Imaging
by: Han, Zihao, et al.
Published: (2025)
by: Han, Zihao, et al.
Published: (2025)
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering
by: Lim, Youngsun, et al.
Published: (2024)
by: Lim, Youngsun, et al.
Published: (2024)
MUST: Modality-Specific Representation-Aware Transformer for Diffusion-Enhanced Survival Prediction with Missing Modality
by: Kim, Kyungwon, et al.
Published: (2026)
by: Kim, Kyungwon, et al.
Published: (2026)
Intermediate Outputs Are More Sensitive Than You Think
by: Huang, Tao, et al.
Published: (2024)
by: Huang, Tao, et al.
Published: (2024)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
When Does Perceptual Alignment Benefit Vision Representations?
by: Sundaram, Shobhita, et al.
Published: (2024)
by: Sundaram, Shobhita, et al.
Published: (2024)
Similar Items
-
DreamCatalyst: Fast and High-Quality 3D Editing via Controlling Editability and Identity Preservation
by: Kim, Jiwook, et al.
Published: (2024) -
Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
by: Yu, Sihyun, et al.
Published: (2024) -
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
by: Feng, Tao, et al.
Published: (2025) -
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026) -
CoT-PL: Chain-of-Thought Pseudo-Labeling for Open-Vocabulary Object Detection
by: Choi, Hojun, et al.
Published: (2025)