Saved in:
| Main Authors: | Kim, Hyunjae, Yoon, Seunghyun, Bui, Trung, Zhao, Handong, Tran, Quan, Dernoncourt, Franck, Kang, Jaewoo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.15120 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025)
by: Zang, Yuan, et al.
Published: (2025)
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024)
by: Argaw, Dawit Mureja, et al.
Published: (2024)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
by: Qi, Daiqing, et al.
Published: (2025)
by: Qi, Daiqing, et al.
Published: (2025)
Supercharged One-step Text-to-Image Diffusion Models with Negative Prompts
by: Nguyen, Viet, et al.
Published: (2024)
by: Nguyen, Viet, et al.
Published: (2024)
Semantic-aware Adversarial Fine-tuning for CLIP
by: Zhang, Jiacheng, et al.
Published: (2026)
by: Zhang, Jiacheng, et al.
Published: (2026)
Benchmarking Direct Preference Optimization for Medical Large Vision-Language Models
by: Kim, Dain, et al.
Published: (2026)
by: Kim, Dain, et al.
Published: (2026)
Robotic-CLIP: Fine-tuning CLIP on Action Data for Robotic Applications
by: Nguyen, Nghia, et al.
Published: (2024)
by: Nguyen, Nghia, et al.
Published: (2024)
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
by: Taesiri, Mohammad Reza, et al.
Published: (2025)
LatteCLIP: Unsupervised CLIP Fine-Tuning via LMM-Synthetic Texts
by: Cao, Anh-Quan, et al.
Published: (2024)
by: Cao, Anh-Quan, et al.
Published: (2024)
CORG: Generating Answers from Complex, Interrelated Contexts
by: Lee, Hyunji, et al.
Published: (2025)
by: Lee, Hyunji, et al.
Published: (2025)
PEEB: Part-based Image Classifiers with an Explainable and Editable Language Bottleneck
by: Pham, Thang M., et al.
Published: (2024)
by: Pham, Thang M., et al.
Published: (2024)
Self-Corrected Flow Distillation for Consistent One-Step and Few-Step Text-to-Image Generation
by: Dao, Quan, et al.
Published: (2024)
by: Dao, Quan, et al.
Published: (2024)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
by: Li, Handong, et al.
Published: (2025)
by: Li, Handong, et al.
Published: (2025)
Fully Fine-tuned CLIP Models are Efficient Few-Shot Learners
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models
by: Morelli, Fabian, et al.
Published: (2026)
by: Morelli, Fabian, et al.
Published: (2026)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
by: Kim, Seoyeon, et al.
Published: (2023)
by: Kim, Seoyeon, et al.
Published: (2023)
Securely Fine-tuning Pre-trained Encoders Against Adversarial Examples
by: Zhou, Ziqi, et al.
Published: (2024)
by: Zhou, Ziqi, et al.
Published: (2024)
SynQ: Accurate Zero-shot Quantization by Synthesis-aware Fine-tuning
by: Kim, Minjun, et al.
Published: (2026)
by: Kim, Minjun, et al.
Published: (2026)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
by: Xiao, Junhao, et al.
Published: (2026)
by: Xiao, Junhao, et al.
Published: (2026)
Realistic Unsupervised CLIP Fine-tuning with Universal Entropy Optimization
by: Liang, Jian, et al.
Published: (2023)
by: Liang, Jian, et al.
Published: (2023)
Prototypical Contrastive Learning-based CLIP Fine-tuning for Object Re-identification
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
Breaking the Limits of Open-Weight CLIP: An Optimization Framework for Self-supervised Fine-tuning of CLIP
by: Mehta, Anant, et al.
Published: (2026)
by: Mehta, Anant, et al.
Published: (2026)
Can LLMs Deceive CLIP? Benchmarking Adversarial Compositionality of Pre-trained Multimodal Representation via Text Updates
by: Ahn, Jaewoo, et al.
Published: (2025)
by: Ahn, Jaewoo, et al.
Published: (2025)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
by: Song, Yehun, et al.
Published: (2025)
by: Song, Yehun, et al.
Published: (2025)
CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Assessing News Thumbnail Representativeness: Counterfactual text can enhance the cross-modal matching ability
by: Yoon, Yejun, et al.
Published: (2024)
by: Yoon, Yejun, et al.
Published: (2024)
DreamBlend: Advancing Personalized Fine-tuning of Text-to-Image Diffusion Models
by: Ram, Shwetha, et al.
Published: (2024)
by: Ram, Shwetha, et al.
Published: (2024)
ReflectCAP: Detailed Image Captioning with Reflective Memory
by: Min, Kyungmin, et al.
Published: (2026)
by: Min, Kyungmin, et al.
Published: (2026)
Enhancing Fine-grained Image Classification through Attentive Batch Training
by: Le, Duy M., et al.
Published: (2024)
by: Le, Duy M., et al.
Published: (2024)
TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection
by: Yoon, Jiae, et al.
Published: (2026)
by: Yoon, Jiae, et al.
Published: (2026)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
by: Rocamora, Elias Abad, et al.
Published: (2025)
by: Rocamora, Elias Abad, et al.
Published: (2025)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
by: Park, NaHyeon, et al.
Published: (2024)
by: Park, NaHyeon, et al.
Published: (2024)
Text2Relight: Creative Portrait Relighting with Text Guidance
by: Cha, Junuk, et al.
Published: (2024)
by: Cha, Junuk, et al.
Published: (2024)
Perturb and Recover: Fine-tuning for Effective Backdoor Removal from CLIP
by: Singh, Naman Deep, et al.
Published: (2024)
by: Singh, Naman Deep, et al.
Published: (2024)
Similar Items
-
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
by: Li, Yifan, et al.
Published: (2026) -
MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos
by: Zang, Yuan, et al.
Published: (2025) -
Scaling Up Video Summarization Pretraining with Large Language Models
by: Argaw, Dawit Mureja, et al.
Published: (2024) -
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025) -
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)