Contextualized Visual Personalization in Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Oh, Yeongtak, Yu, Sangwon, Park, Junsung, Moon, Han Cheol, Mok, Jisoo, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
von: Park, Junsung, et al.
Veröffentlicht: (2025)
von: Park, Junsung, et al.
Veröffentlicht: (2025)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
von: Park, Sangha, et al.
Veröffentlicht: (2025)
von: Park, Sangha, et al.
Veröffentlicht: (2025)
ControlDreamer: Blending Geometry and Style in Text-to-3D
von: Oh, Yeongtak, et al.
Veröffentlicht: (2023)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2023)
Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)
Style-Friendly SNR Sampler for Style-Driven Generation
von: Choi, Jooyoung, et al.
Veröffentlicht: (2024)
von: Choi, Jooyoung, et al.
Veröffentlicht: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
Entropy is not Enough for Test-Time Adaptation: From the Perspective of Disentangled Factors
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
von: Lee, Jonghyun, et al.
Veröffentlicht: (2024)
LogicQA: Logical Anomaly Detection with Vision Language Model Generated Questions
von: Kwon, Yejin, et al.
Veröffentlicht: (2025)
von: Kwon, Yejin, et al.
Veröffentlicht: (2025)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
von: Jung, Mingi, et al.
Veröffentlicht: (2025)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
von: Yi, Jihun, et al.
Veröffentlicht: (2024)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Improving Geometry in Sparse-View 3DGS via Reprojection-based DoF Separation
von: Kim, Yongsung, et al.
Veröffentlicht: (2024)
von: Kim, Yongsung, et al.
Veröffentlicht: (2024)
3D-Aware Vision-Language Models Fine-Tuning with Geometric Distillation
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
von: Lee, Seonho, et al.
Veröffentlicht: (2025)
AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
von: Sung-Bin, Kim, et al.
Veröffentlicht: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
von: Shin, Chaehun, et al.
Veröffentlicht: (2024)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
von: Hyeon-Woo, Nam, et al.
Veröffentlicht: (2024)
Precision matters: Precision-aware ensemble for weakly supervised semantic segmentation
von: Park, Junsung, et al.
Veröffentlicht: (2024)
von: Park, Junsung, et al.
Veröffentlicht: (2024)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
von: Park, Minjeong, et al.
Veröffentlicht: (2025)
Do Vision-Language Models Understand Visual Persuasiveness?
von: Park, Gyuwon
Veröffentlicht: (2025)
von: Park, Gyuwon
Veröffentlicht: (2025)
Balancing Saliency and Coverage: Semantic Prominence-Aware Budgeting for Visual Token Compression in VLMs
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
von: Lee, Jaehoon, et al.
Veröffentlicht: (2026)
Pre-trained Vision and Language Transformers Are Few-Shot Incremental Learners
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
von: Park, Keon-Hee, et al.
Veröffentlicht: (2024)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
Disentangled Motion Modeling for Video Frame Interpolation
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
von: Lew, Jaihyun, et al.
Veröffentlicht: (2024)
Auxiliary Descriptive Knowledge for Few-Shot Adaptation of Vision-Language Model
von: Lee, SuBeen, et al.
Veröffentlicht: (2025)
von: Lee, SuBeen, et al.
Veröffentlicht: (2025)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
von: Lee, Saehyung, et al.
Veröffentlicht: (2024)
Grounding Driving VLA via Inverse Kinematics
von: Park, Junsung, et al.
Veröffentlicht: (2026)
von: Park, Junsung, et al.
Veröffentlicht: (2026)
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
Diffusion-Based Image-to-Image Translation by Noise Correction via Prompt Interpolation
von: Lee, Junsung, et al.
Veröffentlicht: (2024)
von: Lee, Junsung, et al.
Veröffentlicht: (2024)
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2025)
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2025)
SUPER-AD: Semantic Uncertainty-aware Planning for End-to-End Robust Autonomous Driving
von: Ryu, Wonjeong, et al.
Veröffentlicht: (2025)
von: Ryu, Wonjeong, et al.
Veröffentlicht: (2025)
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
von: Song, Jaewoo, et al.
Veröffentlicht: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
von: Demidov, Dmitry, et al.
Veröffentlicht: (2025)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
von: Oh, Yeongtak, et al.
Veröffentlicht: (2025) -
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
von: Lee, Saehyung, et al.
Veröffentlicht: (2024) -
Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
von: Oh, Yeongtak, et al.
Veröffentlicht: (2026) -
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
von: Park, Sangha, et al.
Veröffentlicht: (2025) -
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
von: Oh, Yeongtak, et al.
Veröffentlicht: (2024)