Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tu, Yunbin, Li, Liang, Su, Li, Yan, Chenggang, Huang, Qingming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Context-aware Difference Distilling for Multi-change Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
von: Tao, Zhuo, et al.
Veröffentlicht: (2025)
von: Tao, Zhuo, et al.
Veröffentlicht: (2025)
Enhancing Scientific Figure Captioning Through Cross-modal Learning
von: Rojas, Mateo Alejandro, et al.
Veröffentlicht: (2024)
von: Rojas, Mateo Alejandro, et al.
Veröffentlicht: (2024)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
SOVC: Subject-Oriented Video Captioning
von: Teng, Chang, et al.
Veröffentlicht: (2023)
von: Teng, Chang, et al.
Veröffentlicht: (2023)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
NearID: Identity Representation Learning via Near-identity Distractors
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2026)
von: Cvejic, Aleksandar, et al.
Veröffentlicht: (2026)
Regularized Contrastive Partial Multi-view Outlier Detection
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
von: Wang, Yijia, et al.
Veröffentlicht: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
Lightweight Contrastive Distilled Hashing for Online Cross-modal Retrieval
von: Li, Jiaxing, et al.
Veröffentlicht: (2025)
von: Li, Jiaxing, et al.
Veröffentlicht: (2025)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
von: Ma, Yunchuan, et al.
Veröffentlicht: (2024)
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
von: Ji, Yingrui, et al.
Veröffentlicht: (2025)
von: Ji, Yingrui, et al.
Veröffentlicht: (2025)
C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection
von: Wang, Siheng, et al.
Veröffentlicht: (2025)
von: Wang, Siheng, et al.
Veröffentlicht: (2025)
Contrastive Decoupled Representation Learning and Regularization for Speech-Preserving Facial Expression Manipulation
von: Chen, Tianshui, et al.
Veröffentlicht: (2025)
von: Chen, Tianshui, et al.
Veröffentlicht: (2025)
Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
von: Anonto, Riad Ahmed, et al.
Veröffentlicht: (2025)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
Exploring Structural Degradation in Dense Representations for Self-supervised Learning
von: Dai, Siran, et al.
Veröffentlicht: (2025)
von: Dai, Siran, et al.
Veröffentlicht: (2025)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
Cross-domain EEG-based Emotion Recognition with Contrastive Learning
von: Yan, Rui, et al.
Veröffentlicht: (2025)
von: Yan, Rui, et al.
Veröffentlicht: (2025)
Contrastive Learning with Consistent Representations
von: Wang, Zihu, et al.
Veröffentlicht: (2023)
von: Wang, Zihu, et al.
Veröffentlicht: (2023)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
von: Lv, Henglei, et al.
Veröffentlicht: (2024)
von: Lv, Henglei, et al.
Veröffentlicht: (2024)
Self-supervised Representation Learning with Local Aggregation for Image-based Profiling
von: Dai, Siran, et al.
Veröffentlicht: (2025)
von: Dai, Siran, et al.
Veröffentlicht: (2025)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
SCPNet: Unsupervised Cross-modal Homography Estimation via Intra-modal Self-supervised Learning
von: Zhang, Runmin, et al.
Veröffentlicht: (2024)
von: Zhang, Runmin, et al.
Veröffentlicht: (2024)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
von: Huang, Junming, et al.
Veröffentlicht: (2026)
von: Huang, Junming, et al.
Veröffentlicht: (2026)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
Contrastive Learning for Image Complexity Representation
von: Liu, Shipeng, et al.
Veröffentlicht: (2024)
von: Liu, Shipeng, et al.
Veröffentlicht: (2024)
Representation Alignment Contrastive Regularization for Multi-Object Tracking
von: Liu, Zhonglin, et al.
Veröffentlicht: (2024)
von: Liu, Zhonglin, et al.
Veröffentlicht: (2024)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
LineGS : 3D Line Segment Representation on 3D Gaussian Splatting
von: Yang, Chenggang, et al.
Veröffentlicht: (2024)
von: Yang, Chenggang, et al.
Veröffentlicht: (2024)
Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
von: Li, Bing, et al.
Veröffentlicht: (2022)
von: Li, Bing, et al.
Veröffentlicht: (2022)
Enhancing Perception of Key Changes in Remote Sensing Image Change Captioning
von: Yang, Cong, et al.
Veröffentlicht: (2024)
von: Yang, Cong, et al.
Veröffentlicht: (2024)
When the Future Becomes the Past: Taming Temporal Correspondence for Self-supervised Video Representation Learning
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
CREST: Cross-modal Resonance through Evidential Deep Learning for Enhanced Zero-Shot Learning
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
von: Huang, Haojian, et al.
Veröffentlicht: (2024)
CANeRV: Content Adaptive Neural Representation for Video Compression
von: Tang, Lv, et al.
Veröffentlicht: (2025)
von: Tang, Lv, et al.
Veröffentlicht: (2025)
Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
von: Miao, Peihan, et al.
Veröffentlicht: (2022)
von: Miao, Peihan, et al.
Veröffentlicht: (2022)
Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning
von: Zhang, Yizhi, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhi, et al.
Veröffentlicht: (2025)
Dual-Level Cross-Modal Contrastive Clustering
von: Zhang, Haixin, et al.
Veröffentlicht: (2024)
von: Zhang, Haixin, et al.
Veröffentlicht: (2024)
3D CoCa: Contrastive Learners are 3D Captioners
von: Huang, Ting, et al.
Veröffentlicht: (2025)
von: Huang, Ting, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Context-aware Difference Distilling for Multi-change Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024) -
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024) -
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
von: Tao, Zhuo, et al.
Veröffentlicht: (2025) -
Enhancing Scientific Figure Captioning Through Cross-modal Learning
von: Rojas, Mateo Alejandro, et al.
Veröffentlicht: (2024) -
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)