Context-aware Difference Distilling for Multi-change Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Tu, Yunbin, Li, Liang, Su, Li, Zha, Zheng-Jun, Yan, Chenggang, Huang, Qingming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
di: Tao, Zhuo, et al.
Pubblicazione: (2025)
di: Tao, Zhuo, et al.
Pubblicazione: (2025)
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
di: Pan, Jiadong, et al.
Pubblicazione: (2024)
di: Pan, Jiadong, et al.
Pubblicazione: (2024)
SOVC: Subject-Oriented Video Captioning
di: Teng, Chang, et al.
Pubblicazione: (2023)
di: Teng, Chang, et al.
Pubblicazione: (2023)
Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain Adaptation
di: Deng, Peihua, et al.
Pubblicazione: (2024)
di: Deng, Peihua, et al.
Pubblicazione: (2024)
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
di: Xi, Zeyu, et al.
Pubblicazione: (2024)
di: Xi, Zeyu, et al.
Pubblicazione: (2024)
OPCap:Object-aware Prompting Captioning
di: Huang, Feiyang
Pubblicazione: (2024)
di: Huang, Feiyang
Pubblicazione: (2024)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
di: Yu, Ting, et al.
Pubblicazione: (2024)
di: Yu, Ting, et al.
Pubblicazione: (2024)
Distribution-aware Dataset Distillation for Efficient Image Restoration
di: Zheng, Zhuoran, et al.
Pubblicazione: (2025)
di: Zheng, Zhuoran, et al.
Pubblicazione: (2025)
ViTOC: Vision Transformer and Object-aware Captioner
di: Huang, Feiyang
Pubblicazione: (2024)
di: Huang, Feiyang
Pubblicazione: (2024)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
di: Ma, Yunchuan, et al.
Pubblicazione: (2024)
di: Ma, Yunchuan, et al.
Pubblicazione: (2024)
Leverage Task Context for Object Affordance Ranking
di: Huang, Haojie, et al.
Pubblicazione: (2024)
di: Huang, Haojie, et al.
Pubblicazione: (2024)
SAMKD: Spatial-aware Adaptive Masking Knowledge Distillation for Object Detection
di: Zhang, Zhourui, et al.
Pubblicazione: (2025)
di: Zhang, Zhourui, et al.
Pubblicazione: (2025)
Less is More: Token Context-aware Learning for Object Tracking
di: Xu, Chenlong, et al.
Pubblicazione: (2025)
di: Xu, Chenlong, et al.
Pubblicazione: (2025)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
VIXEN: Visual Text Comparison Network for Image Difference Captioning
di: Black, Alexander, et al.
Pubblicazione: (2024)
di: Black, Alexander, et al.
Pubblicazione: (2024)
Pick-and-Draw: Training-free Semantic Guidance for Text-to-Image Personalization
di: Lv, Henglei, et al.
Pubblicazione: (2024)
di: Lv, Henglei, et al.
Pubblicazione: (2024)
ViDiC: Video Difference Captioning
di: Wu, Jiangtao, et al.
Pubblicazione: (2025)
di: Wu, Jiangtao, et al.
Pubblicazione: (2025)
CLoCKDistill: Consistent Location-and-Context-aware Knowledge Distillation for DETRs
di: Lan, Qizhen, et al.
Pubblicazione: (2025)
di: Lan, Qizhen, et al.
Pubblicazione: (2025)
Uncertainty-aware Long-tailed Weights Model the Utility of Pseudo-labels for Semi-supervised Learning
di: Wu, Jiaqi, et al.
Pubblicazione: (2025)
di: Wu, Jiaqi, et al.
Pubblicazione: (2025)
Frequency-aware Feature Fusion for Dense Image Prediction
di: Chen, Linwei, et al.
Pubblicazione: (2024)
di: Chen, Linwei, et al.
Pubblicazione: (2024)
Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning
di: Lu, Fan, et al.
Pubblicazione: (2024)
di: Lu, Fan, et al.
Pubblicazione: (2024)
Aerial View River Landform Video segmentation: A Weakly Supervised Context-aware Temporal Consistency Distillation Approach
di: Chen, Chi-Han, et al.
Pubblicazione: (2025)
di: Chen, Chi-Han, et al.
Pubblicazione: (2025)
LocCa: Visual Pretraining with Location-aware Captioners
di: Wan, Bo, et al.
Pubblicazione: (2024)
di: Wan, Bo, et al.
Pubblicazione: (2024)
Multi-Exposure Image Fusion via Distilled 3D LUT Grid with Editable Mode
di: Su, Xin, et al.
Pubblicazione: (2024)
di: Su, Xin, et al.
Pubblicazione: (2024)
Knowledge Completes the Vision: A Multimodal Entity-aware Retrieval-Augmented Generation Framework for News Image Captioning
di: You, Xiaoxing, et al.
Pubblicazione: (2025)
di: You, Xiaoxing, et al.
Pubblicazione: (2025)
Knowledge Distillation via the Target-aware Transformer
di: Lin, Sihao, et al.
Pubblicazione: (2022)
di: Lin, Sihao, et al.
Pubblicazione: (2022)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
di: Zheng, Guangcong, et al.
Pubblicazione: (2025)
di: Zheng, Guangcong, et al.
Pubblicazione: (2025)
TrackletGait: A Robust Framework for Gait Recognition in the Wild
di: Zhang, Shaoxiong, et al.
Pubblicazione: (2025)
di: Zhang, Shaoxiong, et al.
Pubblicazione: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
di: Huang, Zhipeng, et al.
Pubblicazione: (2024)
FeatDistill: A Feature Distillation Enhanced Multi-Expert Ensemble Framework for Robust AI-generated Image Detection
di: Tu, Zhilin, et al.
Pubblicazione: (2026)
di: Tu, Zhilin, et al.
Pubblicazione: (2026)
Progressive Homeostatic and Plastic Prompt Tuning for Audio-Visual Multi-Task Incremental Learning
di: Yin, Jiong, et al.
Pubblicazione: (2025)
di: Yin, Jiong, et al.
Pubblicazione: (2025)
A Unified Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
di: Zhu, Jie, et al.
Pubblicazione: (2024)
di: Zhu, Jie, et al.
Pubblicazione: (2024)
Robust Tracking via Mamba-based Context-aware Token Learning
di: Xie, Jinxia, et al.
Pubblicazione: (2024)
di: Xie, Jinxia, et al.
Pubblicazione: (2024)
COLA: Context-aware Language-driven Test-time Adaptation
di: Zhang, Aiming, et al.
Pubblicazione: (2025)
di: Zhang, Aiming, et al.
Pubblicazione: (2025)
Efficient Temporal Sentence Grounding in Videos with Multi-Teacher Knowledge Distillation
di: Liang, Renjie, et al.
Pubblicazione: (2023)
di: Liang, Renjie, et al.
Pubblicazione: (2023)
Exploring Diverse In-Context Configurations for Image Captioning
di: Yang, Xu, et al.
Pubblicazione: (2023)
di: Yang, Xu, et al.
Pubblicazione: (2023)
A Unified and Scalable Membership Inference Method for Visual Self-supervised Encoder via Part-aware Capability
di: Zhu, Jie, et al.
Pubblicazione: (2025)
di: Zhu, Jie, et al.
Pubblicazione: (2025)
Progressive Depth Decoupling and Modulating for Flexible Depth Completion
di: Yang, Zhiwen, et al.
Pubblicazione: (2024)
di: Yang, Zhiwen, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024) -
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024) -
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
di: Tao, Zhuo, et al.
Pubblicazione: (2025) -
SafeCFG: Controlling Harmful Features with Dynamic Safe Guidance for Safe Generation
di: Pan, Jiadong, et al.
Pubblicazione: (2024) -
SOVC: Subject-Oriented Video Captioning
di: Teng, Chang, et al.
Pubblicazione: (2023)