SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Ziping, Xu, Furong, Liu, Jian, Yang, Ming, Guo, Qingpei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
von: Kim, Namho, et al.
Veröffentlicht: (2025)
von: Kim, Namho, et al.
Veröffentlicht: (2025)
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023)
FlattenGPT: Depth Compression for Transformer with Layer Flattening
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
von: Xu, Ruihan, et al.
Veröffentlicht: (2026)
CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
von: Ma, Yunsheng, et al.
Veröffentlicht: (2024)
3D CoCa: Contrastive Learners are 3D Captioners
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
InfMasking: Unleashing Synergistic Information by Contrastive Multimodal Interactions
von: Wen, Liangjian, et al.
Veröffentlicht: (2025)
von: Wen, Liangjian, et al.
Veröffentlicht: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
von: Jiang, Chen, et al.
Veröffentlicht: (2023)
CoCa-CXR: Contrastive Captioners Learn Strong Temporal Structures for Chest X-Ray Vision-Language Understanding
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
von: Chen, Yixiong, et al.
Veröffentlicht: (2025)
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
von: Chen, Tongfei, et al.
Veröffentlicht: (2026)
von: Chen, Tongfei, et al.
Veröffentlicht: (2026)
The More, the Merrier: Contrastive Fusion for Higher-Order Multimodal Alignment
von: Koutoupis, Stefanos, et al.
Veröffentlicht: (2025)
von: Koutoupis, Stefanos, et al.
Veröffentlicht: (2025)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Christos, et al.
Veröffentlicht: (2024)
Towards Multimodal Video Paragraph Captioning Models Robust to Missing Modality
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
von: Chen, Sishuo, et al.
Veröffentlicht: (2024)
BiDepth: A Bidirectional-Depth Neural Network for Spatio-Temporal Prediction
von: Ehsani, Sina, et al.
Veröffentlicht: (2025)
von: Ehsani, Sina, et al.
Veröffentlicht: (2025)
Towards Interpretable Geo-localization: a Concept-Aware Global Image-GPS Alignment Framework
von: Jia, Furong, et al.
Veröffentlicht: (2025)
von: Jia, Furong, et al.
Veröffentlicht: (2025)
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning
von: Zhou, Qifeng, et al.
Veröffentlicht: (2024)
von: Zhou, Qifeng, et al.
Veröffentlicht: (2024)
Symmetric masking strategy enhances the performance of Masked Image Modeling
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh-Binh, et al.
Veröffentlicht: (2024)
QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2025)
von: Kao, Kuei-Chun, et al.
Veröffentlicht: (2025)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
von: Ouyang, Ziheng, et al.
Veröffentlicht: (2025)
STARFlow: Spatial Temporal Feature Re-embedding with Attentive Learning for Real-world Scene Flow
von: Lu, Zhiyang, et al.
Veröffentlicht: (2024)
von: Lu, Zhiyang, et al.
Veröffentlicht: (2024)
Image Captioning in news report scenario
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
von: Liu, Tianrui, et al.
Veröffentlicht: (2024)
Contrast-Prior Enhanced Duality for Mask-Free Shadow Removal
von: Wu, Jiyu, et al.
Veröffentlicht: (2025)
von: Wu, Jiyu, et al.
Veröffentlicht: (2025)
Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
von: Liang, Yongyuan, et al.
Veröffentlicht: (2025)
Top-Down Semantic Refinement for Image Captioning
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
von: Feng, Huidong, et al.
Veröffentlicht: (2026)
von: Feng, Huidong, et al.
Veröffentlicht: (2026)
HOTVCOM: Generating Buzzworthy Comments for Videos
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
von: Chen, Yuyan, et al.
Veröffentlicht: (2024)
Accurate and Scalable Multimodal Pathology Retrieval via Attentive Vision-Language Alignment
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
von: Wang, Hongyi, et al.
Veröffentlicht: (2025)
Multimodal Prompt Alignment for Facial Expression Recognition
von: Ma, Fuyan, et al.
Veröffentlicht: (2025)
von: Ma, Fuyan, et al.
Veröffentlicht: (2025)
Target-Dependent Multimodal Sentiment Analysis Via Employing Visual-to Emotional-Caption Translation Network using Visual-Caption Pairs
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
von: Pandey, Ananya, et al.
Veröffentlicht: (2024)
EntityCLIP: Entity-Centric Image-Text Matching via Multimodal Attentive Contrastive Learning
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
von: Wang, Yaxiong, et al.
Veröffentlicht: (2024)
MaskCD: Mitigating LVLM Hallucinations by Image Head Masked Contrastive Decoding
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
von: Deng, Jingyuan, et al.
Veröffentlicht: (2025)
Shadow Generation with Decomposed Mask Prediction and Attentive Shadow Filling
von: Tao, Xinhao, et al.
Veröffentlicht: (2023)
von: Tao, Xinhao, et al.
Veröffentlicht: (2023)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
von: Zhang, Jusheng, et al.
Veröffentlicht: (2025)
Describe Anything: Detailed Localized Image and Video Captioning
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Alternative Telescopic Displacement: An Efficient Multimodal Alignment Method
von: Qin, Jiahao, et al.
Veröffentlicht: (2023)
von: Qin, Jiahao, et al.
Veröffentlicht: (2023)
QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
von: Wang, Jiale, et al.
Veröffentlicht: (2026)
GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures
von: Patock, Jake R., et al.
Veröffentlicht: (2025)
von: Patock, Jake R., et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024) -
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
von: Kim, Namho, et al.
Veröffentlicht: (2025) -
Pink: Unveiling the Power of Referential Comprehension for Multi-modal LLMs
von: Xuan, Shiyu, et al.
Veröffentlicht: (2023) -
FlattenGPT: Depth Compression for Transformer with Layer Flattening
von: Xu, Ruihan, et al.
Veröffentlicht: (2026) -
CSMCIR: CoT-Enhanced Symmetric Alignment with Memory Bank for Composed Image Retrieval
von: Qian, Zhipeng, et al.
Veröffentlicht: (2026)