Saved in:
| Main Authors: | Zhang, Zhiwei, Liu, Yuliang |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2303.05983 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026)
by: Wu, Junfei, et al.
Published: (2026)
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
by: Cai, Yuliang, et al.
Published: (2024)
by: Cai, Yuliang, et al.
Published: (2024)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
by: Liu, Zichuan, et al.
Published: (2025)
by: Liu, Zichuan, et al.
Published: (2025)
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
by: Feng, Zhanbo, et al.
Published: (2023)
by: Feng, Zhanbo, et al.
Published: (2023)
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)
by: Xie, Chunyu, et al.
Published: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
START: Spatial and Textual Learning for Chart Understanding
by: Liu, Zhuoming, et al.
Published: (2025)
by: Liu, Zhuoming, et al.
Published: (2025)
CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions
by: Zhan, Yuliang, et al.
Published: (2026)
by: Zhan, Yuliang, et al.
Published: (2026)
Contact-aware Human Motion Generation from Textual Descriptions
by: Ma, Sihan, et al.
Published: (2024)
by: Ma, Sihan, et al.
Published: (2024)
Unlocking Attributes' Contribution to Successful Camouflage: A Combined Textual and VisualAnalysis Strategy
by: Zhang, Hong, et al.
Published: (2024)
by: Zhang, Hong, et al.
Published: (2024)
Enhancing Spatial Reasoning through Visual and Textual Thinking
by: Liang, Xun, et al.
Published: (2025)
by: Liang, Xun, et al.
Published: (2025)
GPTDrawer: Enhancing Visual Synthesis through ChatGPT
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
CLIP Model for Images to Textual Prompts Based on Top-k Neighbors
by: Zhang, Xin, et al.
Published: (2024)
by: Zhang, Xin, et al.
Published: (2024)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
by: Cornet, Clément, et al.
Published: (2025)
by: Cornet, Clément, et al.
Published: (2025)
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy
by: Yang, Te, et al.
Published: (2024)
by: Yang, Te, et al.
Published: (2024)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
by: Zhang, Beichen, et al.
Published: (2025)
by: Zhang, Beichen, et al.
Published: (2025)
Human-Object Interaction from Human-Level Instructions
by: Wu, Zhen, et al.
Published: (2024)
by: Wu, Zhen, et al.
Published: (2024)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
by: An, Yanru, et al.
Published: (2025)
by: An, Yanru, et al.
Published: (2025)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
ViscoNet: Bridging and Harmonizing Visual and Textual Conditioning for ControlNet
by: Cheong, Soon Yau, et al.
Published: (2023)
by: Cheong, Soon Yau, et al.
Published: (2023)
VVTRec: Radio Interferometric Reconstruction through Visual and Textual Modality Enrichment
by: Cheng, Kai, et al.
Published: (2026)
by: Cheng, Kai, et al.
Published: (2026)
Synthesize Privacy-Preserving High-Resolution Images via Private Textual Intermediaries
by: Wang, Haoxiang, et al.
Published: (2025)
by: Wang, Haoxiang, et al.
Published: (2025)
Human Motion Instruction Tuning
by: Li, Lei, et al.
Published: (2024)
by: Li, Lei, et al.
Published: (2024)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
by: Yang, Tianyu, et al.
Published: (2024)
by: Yang, Tianyu, et al.
Published: (2024)
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
by: Zhang, Yiman, et al.
Published: (2025)
by: Zhang, Yiman, et al.
Published: (2025)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
ChatBEV: A Visual Language Model that Understands BEV Maps
by: Xu, Qingyao, et al.
Published: (2025)
by: Xu, Qingyao, et al.
Published: (2025)
Beyond Images: Adaptive Fusion of Visual and Textual Data for Food Classification
by: Mittal, Prateek, et al.
Published: (2023)
by: Mittal, Prateek, et al.
Published: (2023)
ChatGen: Automatic Text-to-Image Generation From FreeStyle Chatting
by: Jia, Chengyou, et al.
Published: (2024)
by: Jia, Chengyou, et al.
Published: (2024)
VIGC: Visual Instruction Generation and Correction
by: Wang, Bin, et al.
Published: (2023)
by: Wang, Bin, et al.
Published: (2023)
Breaking Language Barriers in Visual Language Models via Multilingual Textual Regularization
by: Pikabea, Iñigo, et al.
Published: (2025)
by: Pikabea, Iñigo, et al.
Published: (2025)
Hierarchical Prompt Learning for Image- and Text-Based Person Re-Identification
by: Zhou, Linhan, et al.
Published: (2025)
by: Zhou, Linhan, et al.
Published: (2025)
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
by: Wu, Mingrui, et al.
Published: (2025)
by: Wu, Mingrui, et al.
Published: (2025)
Generalizable Object Re-Identification via Visual In-Context Prompting
by: Huang, Zhizhong, et al.
Published: (2025)
by: Huang, Zhizhong, et al.
Published: (2025)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
by: Mishra, Abhijit, et al.
Published: (2025)
by: Mishra, Abhijit, et al.
Published: (2025)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
by: Yang, Xiaobo, et al.
Published: (2025)
by: Yang, Xiaobo, et al.
Published: (2025)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
by: Yin, Xiangyu, et al.
Published: (2025)
by: Yin, Xiangyu, et al.
Published: (2025)
PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
by: Liu, Jiazhen, et al.
Published: (2024)
by: Liu, Jiazhen, et al.
Published: (2024)
Similar Items
-
Chatting with Images for Introspective Visual Thinking
by: Wu, Junfei, et al.
Published: (2026) -
CluMo: Cluster-based Modality Fusion Prompt for Continual Learning in Visual Question Answering
by: Cai, Yuliang, et al.
Published: (2024) -
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
by: Liu, Zichuan, et al.
Published: (2025) -
Textual and Visual Prompt Fusion for Image Editing via Step-Wise Alignment
by: Feng, Zhanbo, et al.
Published: (2023) -
FG-CLIP: Fine-Grained Visual and Textual Alignment
by: Xie, Chunyu, et al.
Published: (2025)