Language-Image Alignment with Fixed Text Encoders
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Jingfeng, Wu, Ziyang, Zhao, Yue, Ma, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UNIT: Unifying Image and Text Recognition in One Vision Encoder
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
von: Bachu, Saketh, et al.
Veröffentlicht: (2024)
von: Bachu, Saketh, et al.
Veröffentlicht: (2024)
Text-Guided Semantic Image Encoder
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2025)
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2025)
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
von: Wu, Jiong, et al.
Veröffentlicht: (2025)
von: Wu, Jiong, et al.
Veröffentlicht: (2025)
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
von: Kou, Siqi, et al.
Veröffentlicht: (2026)
von: Kou, Siqi, et al.
Veröffentlicht: (2026)
TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment
von: Liu, Jiarun, et al.
Veröffentlicht: (2026)
von: Liu, Jiarun, et al.
Veröffentlicht: (2026)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
von: Park, NaHyeon, et al.
Veröffentlicht: (2024)
Video-Panda: Parameter-efficient Alignment for Encoder-free Video-Language Models
von: Yi, Jinhui, et al.
Veröffentlicht: (2024)
von: Yi, Jinhui, et al.
Veröffentlicht: (2024)
Scaling Down Text Encoders of Text-to-Image Diffusion Models
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
von: Wang, Lifu, et al.
Veröffentlicht: (2025)
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
von: Zhou, Dewei, et al.
Veröffentlicht: (2025)
Image Generation with a Sphere Encoder
von: Yue, Kaiyu, et al.
Veröffentlicht: (2026)
von: Yue, Kaiyu, et al.
Veröffentlicht: (2026)
TAMISeg: Text-Aligned Multi-scale Medical Image Segmentation with Semantic Encoder Distillation
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
von: Gao, Qiang, et al.
Veröffentlicht: (2026)
InstructEngine: Instruction-driven Text-to-Image Alignment
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
Image Regeneration: Evaluating Text-to-Image Model via Generating Identical Image with Multimodal Large Language Models
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
von: Meng, Chutian, et al.
Veröffentlicht: (2024)
Day-Night Adaptation: An Innovative Source-free Adaptation Framework for Medical Image Segmentation
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
How to Make Cross Encoder a Good Teacher for Efficient Image-Text Retrieval?
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
von: Chen, Yuxin, et al.
Veröffentlicht: (2024)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
LCM-Lookahead for Encoder-based Text-to-Image Personalization
von: Gal, Rinon, et al.
Veröffentlicht: (2024)
von: Gal, Rinon, et al.
Veröffentlicht: (2024)
FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
von: Pan, Kaihang, et al.
Veröffentlicht: (2025)
Beyond Fixed Inference: Quantitative Flow Matching for Adaptive Image Denoising
von: Duan, Jigang, et al.
Veröffentlicht: (2026)
von: Duan, Jigang, et al.
Veröffentlicht: (2026)
Restoring Initial Noise Sensitivity in Text-to-Image Distillation via Geometric Alignment
von: Huang, Huayang, et al.
Veröffentlicht: (2026)
von: Huang, Huayang, et al.
Veröffentlicht: (2026)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
von: Li, Senmao, et al.
Veröffentlicht: (2025)
von: Li, Senmao, et al.
Veröffentlicht: (2025)
TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
von: Zhao, Zixu, et al.
Veröffentlicht: (2025)
UniSemAlign: Text-Prototype Alignment with a Foundation Encoder for Semi-Supervised Histopathology Segmentation
von: Thai, Le-Van, et al.
Veröffentlicht: (2026)
von: Thai, Le-Van, et al.
Veröffentlicht: (2026)
CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
von: Ma, Lichen, et al.
Veröffentlicht: (2024)
Fast Prompt Alignment for Text-to-Image Generation
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
von: Mrini, Khalil, et al.
Veröffentlicht: (2024)
Gradient Alignment Improves Test-Time Adaptation for Medical Image Segmentation
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
Hierarchical Vision-Language Alignment for Text-to-Image Generation via Diffusion Models
von: Johnson, Emily, et al.
Veröffentlicht: (2025)
von: Johnson, Emily, et al.
Veröffentlicht: (2025)
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
von: Xing, Xiaoying, et al.
Veröffentlicht: (2025)
Graph Domain Adaptation with Dual-branch Encoder and Two-level Alignment for Whole Slide Image-based Survival Prediction
von: Shou, Yuntao, et al.
Veröffentlicht: (2024)
von: Shou, Yuntao, et al.
Veröffentlicht: (2024)
Enhancing Diffusion Models with Text-Encoder Reinforcement Learning
von: Chen, Chaofeng, et al.
Veröffentlicht: (2023)
von: Chen, Chaofeng, et al.
Veröffentlicht: (2023)
Detecting Deepfakes with Multivariate Soft Blending and CLIP-based Image-Text Alignment
von: Li, Jingwei, et al.
Veröffentlicht: (2026)
von: Li, Jingwei, et al.
Veröffentlicht: (2026)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
von: Li, Handong, et al.
Veröffentlicht: (2025)
von: Li, Handong, et al.
Veröffentlicht: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
von: Kim, Seoyeon, et al.
Veröffentlicht: (2023)
von: Kim, Seoyeon, et al.
Veröffentlicht: (2023)
Object Fidelity Diffusion for Remote Sensing Image Generation
von: Ye, Ziqi, et al.
Veröffentlicht: (2025)
von: Ye, Ziqi, et al.
Veröffentlicht: (2025)
MIGC: Multi-Instance Generation Controller for Text-to-Image Synthesis
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
von: Zhou, Dewei, et al.
Veröffentlicht: (2024)
TAI++: Text as Image for Multi-Label Image Classification by Co-Learning Transferable Prompt
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
von: Wu, Xiangyu, et al.
Veröffentlicht: (2024)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
von: Wang, Juntong, et al.
Veröffentlicht: (2025)
Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
UNIT: Unifying Image and Text Recognition in One Vision Encoder
von: Zhu, Yi, et al.
Veröffentlicht: (2024) -
Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language Models
von: Bachu, Saketh, et al.
Veröffentlicht: (2024) -
Text-Guided Semantic Image Encoder
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2025) -
CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation
von: Wu, Jiong, et al.
Veröffentlicht: (2025) -
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
von: Kou, Siqi, et al.
Veröffentlicht: (2026)