An Intermediate Fusion ViT Enables Efficient Text-Image Alignment in Diffusion Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Zizhao, Jia, Shaochong, Rostami, Mohammad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
von: Hu, Youbing, et al.
Veröffentlicht: (2024)
von: Hu, Youbing, et al.
Veröffentlicht: (2024)
Hybrid CNN-ViT Framework for Motion-Blurred Scene Text Restoration
von: Rashid, Umar, et al.
Veröffentlicht: (2025)
von: Rashid, Umar, et al.
Veröffentlicht: (2025)
Lateralization MLP: A Simple Brain-inspired Architecture for Diffusion
von: Hu, Zizhao, et al.
Veröffentlicht: (2024)
von: Hu, Zizhao, et al.
Veröffentlicht: (2024)
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025)
von: Shah, Arya, et al.
Veröffentlicht: (2025)
DiffPoint: Single and Multi-view Point Cloud Reconstruction with ViT Based Diffusion Model
von: Feng, Yu, et al.
Veröffentlicht: (2024)
von: Feng, Yu, et al.
Veröffentlicht: (2024)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
von: Xu, Xuwei, et al.
Veröffentlicht: (2023)
von: Xu, Xuwei, et al.
Veröffentlicht: (2023)
SAC-ViT: Semantic-Aware Clustering Vision Transformer with Early Exit
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
PEANO-ViT: Power-Efficient Approximations of Non-Linearities in Vision Transformers
von: Sadeghi, Mohammad Erfan, et al.
Veröffentlicht: (2024)
von: Sadeghi, Mohammad Erfan, et al.
Veröffentlicht: (2024)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
von: Cai, Yuliang, et al.
Veröffentlicht: (2024)
von: Cai, Yuliang, et al.
Veröffentlicht: (2024)
Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision Transformer
von: Shi, Huihong, et al.
Veröffentlicht: (2024)
von: Shi, Huihong, et al.
Veröffentlicht: (2024)
SFMViT: SlowFast Meet ViT in Chaotic World
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
DFQ-ViT: Data-Free Quantization for Vision Transformers without Fine-tuning
von: Tong, Yujia, et al.
Veröffentlicht: (2025)
von: Tong, Yujia, et al.
Veröffentlicht: (2025)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
von: Xian, Jia Jun Cheng, et al.
Veröffentlicht: (2025)
von: Xian, Jia Jun Cheng, et al.
Veröffentlicht: (2025)
HydraViT: Stacking Heads for a Scalable ViT
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
von: Haberer, Janek, et al.
Veröffentlicht: (2024)
Personalized Safety Alignment for Text-to-Image Diffusion Models
von: Lei, Yu, et al.
Veröffentlicht: (2025)
von: Lei, Yu, et al.
Veröffentlicht: (2025)
Instant Preference Alignment for Text-to-Image Diffusion Models
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
Tiny-ViT: A Compact Vision Transformer for Efficient and Explainable Potato Leaf Disease Classification
von: Mia, Shakil, et al.
Veröffentlicht: (2026)
von: Mia, Shakil, et al.
Veröffentlicht: (2026)
Sub-token ViT Embedding via Stochastic Resonance Transformers
von: Lao, Dong, et al.
Veröffentlicht: (2023)
von: Lao, Dong, et al.
Veröffentlicht: (2023)
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation
von: Li, Siyou, et al.
Veröffentlicht: (2024)
von: Li, Siyou, et al.
Veröffentlicht: (2024)
Case-Enhanced Vision Transformer: Improving Explanations of Image Similarity with a ViT-based Similarity Metric
von: Zhao, Ziwei, et al.
Veröffentlicht: (2024)
von: Zhao, Ziwei, et al.
Veröffentlicht: (2024)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
Knowledge Distillation in YOLOX-ViT for Side-Scan Sonar Object Detection
von: Aubard, Martin, et al.
Veröffentlicht: (2024)
von: Aubard, Martin, et al.
Veröffentlicht: (2024)
MobilePlantViT: A Mobile-friendly Hybrid ViT for Generalized Plant Disease Image Classification
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
von: Tonmoy, Moshiur Rahman, et al.
Veröffentlicht: (2025)
Uncovering the Text Embedding in Text-to-Image Diffusion Models
von: Yu, Hu, et al.
Veröffentlicht: (2024)
von: Yu, Hu, et al.
Veröffentlicht: (2024)
Communication Efficient Split Learning of ViTs with Attention-based Double Compression
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
von: Alvetreti, Federico, et al.
Veröffentlicht: (2025)
Filtered-ViT: A Robust Defense Against Multiple Adversarial Patch Attacks
von: Khanal, Aja, et al.
Veröffentlicht: (2025)
von: Khanal, Aja, et al.
Veröffentlicht: (2025)
VIVID-Med: LLM-Supervised Structured Pretraining for Deployable Medical ViTs
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
von: Wang, Xiyao, et al.
Veröffentlicht: (2026)
Curvature Diversity-Driven Deformation and Domain Alignment for Point Cloud
von: Wu, Mengxi, et al.
Veröffentlicht: (2024)
von: Wu, Mengxi, et al.
Veröffentlicht: (2024)
CNN-ViT Fusion with Adaptive Attention Gate for Brain Tumor MRI Classification: A Hybrid Deep Learning Model
von: Hasnain, Syed Ibad, et al.
Veröffentlicht: (2026)
von: Hasnain, Syed Ibad, et al.
Veröffentlicht: (2026)
InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
von: Rastegar, Sarah, et al.
Veröffentlicht: (2025)
von: Rastegar, Sarah, et al.
Veröffentlicht: (2025)
Derm-T2IM: Harnessing Synthetic Skin Lesion Data via Stable Diffusion Models for Enhanced Skin Disease Classification using ViT and CNN
von: Farooq, Muhammad Ali, et al.
Veröffentlicht: (2024)
von: Farooq, Muhammad Ali, et al.
Veröffentlicht: (2024)
Implicit to Explicit Entropy Regularization: Benchmarking ViT Fine-tuning under Noisy Labels
von: Marrium, Maria, et al.
Veröffentlicht: (2024)
von: Marrium, Maria, et al.
Veröffentlicht: (2024)
ViTs are Everywhere: A Comprehensive Study Showcasing Vision Transformers in Different Domain
von: Mia, Md Sohag, et al.
Veröffentlicht: (2023)
von: Mia, Md Sohag, et al.
Veröffentlicht: (2023)
ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2025)
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2025)
H-CNN-ViT: A Hierarchical Gated Attention Multi-Branch Model for Bladder Cancer Recurrence Prediction
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
von: Li, Xueyang, et al.
Veröffentlicht: (2025)
ViFusionTST: Deep Fusion of Time-Series Image Representations from Load Signals for Early Bed-Exit Prediction
von: Liu, Hao, et al.
Veröffentlicht: (2025)
von: Liu, Hao, et al.
Veröffentlicht: (2025)
Alignment-Guided Score Matching for Text-to-Image Alignment in Diffusion Models
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
von: Lee, Jaa-Yeon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition
von: Hu, Youbing, et al.
Veröffentlicht: (2024) -
Hybrid CNN-ViT Framework for Motion-Blurred Scene Text Restoration
von: Rashid, Umar, et al.
Veröffentlicht: (2025) -
Lateralization MLP: A Simple Brain-inspired Architecture for Diffusion
von: Hu, Zizhao, et al.
Veröffentlicht: (2024) -
Purrturbed but Stable: Human-Cat Invariant Representations Across CNNs, ViTs and Self-Supervised ViTs
von: Shah, Arya, et al.
Veröffentlicht: (2025) -
DiffPoint: Single and Multi-view Point Cloud Reconstruction with ViT Based Diffusion Model
von: Feng, Yu, et al.
Veröffentlicht: (2024)