TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Yao, Fu, Qiqian, Tao, Heyi, Wu, Yuqun, Zhu, Zhen, Hoiem, Derek |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023)
by: Zhu, Zhen, et al.
Published: (2023)
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025)
by: Sahoo, Aadarsh, et al.
Published: (2025)
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
by: Lee, Jae Yong, et al.
Published: (2024)
by: Lee, Jae Yong, et al.
Published: (2024)
AttnDreamBooth: Towards Text-Aligned Personalized Text-to-Image Generation
by: Pang, Lianyu, et al.
Published: (2024)
by: Pang, Lianyu, et al.
Published: (2024)
Foley Control: Aligning a Frozen Latent Text-to-Audio Model to Video
by: Rowles, Ciara, et al.
Published: (2025)
by: Rowles, Ciara, et al.
Published: (2025)
Text Region Multiple Information Perception Network for Scene Text Detection
by: Zheng, Jinzhi, et al.
Published: (2024)
by: Zheng, Jinzhi, et al.
Published: (2024)
Text-Region Matching for Multi-Label Image Recognition with Missing Labels
by: Ma, Leilei, et al.
Published: (2024)
by: Ma, Leilei, et al.
Published: (2024)
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024)
by: Khosla, Savya, et al.
Published: (2024)
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
by: Hu, Zijing, et al.
Published: (2025)
by: Hu, Zijing, et al.
Published: (2025)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
Anytime Continual Learning for Open Vocabulary Classification
by: Zhu, Zhen, et al.
Published: (2024)
by: Zhu, Zhen, et al.
Published: (2024)
How to Teach Large Multimodal Models New Skills
by: Zhu, Zhen, et al.
Published: (2025)
by: Zhu, Zhen, et al.
Published: (2025)
Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative Grounding
by: Li, Hongyu, et al.
Published: (2024)
by: Li, Hongyu, et al.
Published: (2024)
Bridging the Gap: Aligning Text-to-Image Diffusion Models with Specific Feedback
by: Niu, Xuexiang, et al.
Published: (2024)
by: Niu, Xuexiang, et al.
Published: (2024)
CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
by: Jiang, Dongzhi, et al.
Published: (2024)
by: Jiang, Dongzhi, et al.
Published: (2024)
Fair Text to Medical Image Diffusion Model with Subgroup Distribution Aligned Tuning
by: Han, Xu, et al.
Published: (2024)
by: Han, Xu, et al.
Published: (2024)
Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long Text
by: Liang, Guotao, et al.
Published: (2025)
by: Liang, Guotao, et al.
Published: (2025)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
by: Kim, Dongwon, et al.
Published: (2025)
by: Kim, Dongwon, et al.
Published: (2025)
A Token-level Text Image Foundation Model for Document Understanding
by: Guan, Tongkun, et al.
Published: (2025)
by: Guan, Tongkun, et al.
Published: (2025)
SafeCtrl: Region-Based Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress
by: Zhang, Lingyun, et al.
Published: (2025)
by: Zhang, Lingyun, et al.
Published: (2025)
SceneDiff: A Benchmark and Method for Multiview Object Change Detection
by: Wu, Yuqun, et al.
Published: (2025)
by: Wu, Yuqun, et al.
Published: (2025)
AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models
by: Agarwal, Aishwarya, et al.
Published: (2024)
by: Agarwal, Aishwarya, et al.
Published: (2024)
Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimization
by: Liu, Zhuohan, et al.
Published: (2026)
by: Liu, Zhuohan, et al.
Published: (2026)
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
TokenCompose: Text-to-Image Diffusion with Token-level Supervision
by: Wang, Zirui, et al.
Published: (2023)
by: Wang, Zirui, et al.
Published: (2023)
UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
by: Ma, Lichen, et al.
Published: (2026)
by: Ma, Lichen, et al.
Published: (2026)
SafeText: Safe Text-to-image Models via Aligning the Text Encoder
by: Hu, Yuepeng, et al.
Published: (2025)
by: Hu, Yuepeng, et al.
Published: (2025)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
by: Lin, Xingtao, et al.
Published: (2024)
by: Lin, Xingtao, et al.
Published: (2024)
PI3D: Efficient Text-to-3D Generation with Pseudo-Image Diffusion
by: Liu, Ying-Tian, et al.
Published: (2023)
by: Liu, Ying-Tian, et al.
Published: (2023)
Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement
by: Chen, Zhennan, et al.
Published: (2024)
by: Chen, Zhennan, et al.
Published: (2024)
Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation
by: Xing, Xiaoying, et al.
Published: (2025)
by: Xing, Xiaoying, et al.
Published: (2025)
From Text to Mask: Localizing Entities Using the Attention of Text-to-Image Diffusion Models
by: Xiao, Changming, et al.
Published: (2023)
by: Xiao, Changming, et al.
Published: (2023)
Discriminative Class Tokens for Text-to-Image Diffusion Models
by: Schwartz, Idan, et al.
Published: (2023)
by: Schwartz, Idan, et al.
Published: (2023)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Image2Text2Image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-to-Image Diffusion Models
by: Huang, Jia-Hong, et al.
Published: (2024)
by: Huang, Jia-Hong, et al.
Published: (2024)
Skill-Aligned Annotation for Reliable Evaluation in Text-to-Image Generation
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
by: Eldesokey, Abdelrahman, et al.
Published: (2026)
Similar Items
-
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026) -
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024) -
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023) -
Aligning Text, Images, and 3D Structure Token-by-Token
by: Sahoo, Aadarsh, et al.
Published: (2025) -
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)