TagAlign: Improving Vision-Language Alignment with Multi-Tag Classification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Qinying, Wu, Wei, Zheng, Kecheng, Tong, Zhan, Liu, Jiawei, Liu, Yu, Chen, Wei, Wang, Zilei, Shen, Yujun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024)
von: Lu, Fan, et al.
Veröffentlicht: (2024)
Convex Combination Consistency between Neighbors for Weakly-supervised Action Localization
von: Liu, Qinying, et al.
Veröffentlicht: (2022)
von: Liu, Qinying, et al.
Veröffentlicht: (2022)
Tag2Text: Guiding Vision-Language Model via Image Tagging
von: Huang, Xinyu, et al.
Veröffentlicht: (2023)
von: Huang, Xinyu, et al.
Veröffentlicht: (2023)
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
von: Wang, Hanlin, et al.
Veröffentlicht: (2024)
MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
von: Chen, Siyi, et al.
Veröffentlicht: (2025)
von: Chen, Siyi, et al.
Veröffentlicht: (2025)
Modest-Align: Data-Efficient Alignment for Vision-Language Models
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
Exploring Motion-Language Alignment for Text-driven Motion Generation
von: Gu, Ruxi, et al.
Veröffentlicht: (2026)
von: Gu, Ruxi, et al.
Veröffentlicht: (2026)
ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
MePT: Multi-Representation Guided Prompt Tuning for Vision-Language Model
von: Wang, Xinyang, et al.
Veröffentlicht: (2024)
von: Wang, Xinyang, et al.
Veröffentlicht: (2024)
LoTLIP: Improving Language-Image Pre-training for Long Text Understanding
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
TagCLIP: Improving Discrimination Ability of Open-Vocabulary Semantic Segmentation
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
von: Li, Jingyao, et al.
Veröffentlicht: (2023)
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias
von: Jo, Sanghyun, et al.
Veröffentlicht: (2024)
von: Jo, Sanghyun, et al.
Veröffentlicht: (2024)
DreamLIP: Language-Image Pre-training with Long Captions
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
von: Zheng, Kecheng, et al.
Veröffentlicht: (2024)
TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection
von: Chen, Jiankang, et al.
Veröffentlicht: (2024)
von: Chen, Jiankang, et al.
Veröffentlicht: (2024)
Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs
von: Liu, Shi, et al.
Veröffentlicht: (2024)
von: Liu, Shi, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Model with Unmasked Token Alignment
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
von: Rivera, Antonio Carlos, et al.
Veröffentlicht: (2024)
FolkTalent: Enhancing Classification and Tagging of Indian Folk Paintings
von: Hada, Nancy, et al.
Veröffentlicht: (2024)
von: Hada, Nancy, et al.
Veröffentlicht: (2024)
Tag-Enriched Multi-Attention with Large Language Models for Cross-Domain Sequential Recommendation
von: Wu, Wangyu, et al.
Veröffentlicht: (2025)
von: Wu, Wangyu, et al.
Veröffentlicht: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
von: Yan, Ziang, et al.
Veröffentlicht: (2024)
Correcting Class Imbalances with Self-Training for Improved Universal Lesion Detection and Tagging
von: Shieh, Alexander, et al.
Veröffentlicht: (2025)
von: Shieh, Alexander, et al.
Veröffentlicht: (2025)
OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
von: Ouyang, Jieer, et al.
Veröffentlicht: (2025)
von: Ouyang, Jieer, et al.
Veröffentlicht: (2025)
$Δ\mathrm{Energy}$: Optimizing Energy Change During Vision-Language Alignment Improves both OOD Detection and OOD Generalization
von: Zhu, Lin, et al.
Veröffentlicht: (2025)
von: Zhu, Lin, et al.
Veröffentlicht: (2025)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
von: Qi, Daiqing, et al.
Veröffentlicht: (2024)
Exploring Visual Vulnerabilities via Multi-Loss Adversarial Search for Jailbreaking Vision-Language Models
von: Hao, Shuyang, et al.
Veröffentlicht: (2024)
von: Hao, Shuyang, et al.
Veröffentlicht: (2024)
Are Multimodal Large Language Models Good Annotators for Image Tagging?
von: Xie, Ming-Kun, et al.
Veröffentlicht: (2026)
von: Xie, Ming-Kun, et al.
Veröffentlicht: (2026)
YoloTag: Vision-based Robust UAV Navigation with Fiducial Markers
von: Raxit, Sourav, et al.
Veröffentlicht: (2024)
von: Raxit, Sourav, et al.
Veröffentlicht: (2024)
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
UKnow: A Unified Knowledge Protocol with Multimodal Knowledge Graph Datasets for Reasoning and Vision-Language Pre-Training
von: Gong, Biao, et al.
Veröffentlicht: (2023)
von: Gong, Biao, et al.
Veröffentlicht: (2023)
TagGAN: A Generative Model for Data Tagging
von: Nawaz, Muhammad, et al.
Veröffentlicht: (2025)
von: Nawaz, Muhammad, et al.
Veröffentlicht: (2025)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
von: Wijaya, Robert, et al.
Veröffentlicht: (2024)
von: Wijaya, Robert, et al.
Veröffentlicht: (2024)
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
von: Yang, Xiao, et al.
Veröffentlicht: (2026)
Mimir: Improving Video Diffusion Models for Precise Text Understanding
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
von: Tan, Shuai, et al.
Veröffentlicht: (2024)
TagOOD: A Novel Approach to Out-of-Distribution Detection via Vision-Language Representations and Class Center Learning
von: Li, Jinglun, et al.
Veröffentlicht: (2024)
von: Li, Jinglun, et al.
Veröffentlicht: (2024)
Learning Visual Generative Priors without Text
von: Ma, Shuailei, et al.
Veröffentlicht: (2024)
von: Ma, Shuailei, et al.
Veröffentlicht: (2024)
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
von: Khosla, Savya, et al.
Veröffentlicht: (2026)
von: Khosla, Savya, et al.
Veröffentlicht: (2026)
3D Universal Lesion Detection and Tagging in CT with Self-Training
von: Frazier, Jared, et al.
Veröffentlicht: (2025)
von: Frazier, Jared, et al.
Veröffentlicht: (2025)
CLIP the Landscape: Automated Tagging of Crowdsourced Landscape Images
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2025)
von: Ilyankou, Ilya, et al.
Veröffentlicht: (2025)
Hierarchical Co-Embedding of Font Shapes and Impression Tags
von: Kubota, Yugo, et al.
Veröffentlicht: (2026)
von: Kubota, Yugo, et al.
Veröffentlicht: (2026)
ExpAlign: Expectation-Guided Vision-Language Alignment for Open-Vocabulary Grounding
von: Hu, Junyi, et al.
Veröffentlicht: (2026)
von: Hu, Junyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
von: Lu, Fan, et al.
Veröffentlicht: (2024) -
Convex Combination Consistency between Neighbors for Weakly-supervised Action Localization
von: Liu, Qinying, et al.
Veröffentlicht: (2022) -
Tag2Text: Guiding Vision-Language Model via Image Tagging
von: Huang, Xinyu, et al.
Veröffentlicht: (2023) -
Contextual AD Narration with Interleaved Multimodal Sequence
von: Wang, Hanlin, et al.
Veröffentlicht: (2024) -
MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
von: Chen, Siyi, et al.
Veröffentlicht: (2025)