RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Yunchuan, Qing, Laiyun, Li, Guorong, Qi, Yuankai, Beheshti, Amin, Sheng, Quan Z., Huang, Qingming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SOVC: Subject-Oriented Video Captioning
di: Teng, Chang, et al.
Pubblicazione: (2023)
di: Teng, Chang, et al.
Pubblicazione: (2023)
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
di: Ma, Yunchuan, et al.
Pubblicazione: (2026)
di: Ma, Yunchuan, et al.
Pubblicazione: (2026)
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
di: Ma, Yunchuan, et al.
Pubblicazione: (2026)
di: Ma, Yunchuan, et al.
Pubblicazione: (2026)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
di: Tian, Mingkai, et al.
Pubblicazione: (2025)
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
di: Zhao, Yiming, et al.
Pubblicazione: (2025)
Foresee-to-Ground: From Predictive Temporal Perception to Evidence-Driven Reasoning for Video Temporal Grounding
di: Zheng, Zelin, et al.
Pubblicazione: (2026)
di: Zheng, Zelin, et al.
Pubblicazione: (2026)
FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing
di: Cong, Gaoxiang, et al.
Pubblicazione: (2025)
di: Cong, Gaoxiang, et al.
Pubblicazione: (2025)
Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos
di: Ma, Jianbo, et al.
Pubblicazione: (2025)
di: Ma, Jianbo, et al.
Pubblicazione: (2025)
ProgRoCC: A Progressive Approach to Rough Crowd Counting
di: Jiang, Shengqin, et al.
Pubblicazione: (2025)
di: Jiang, Shengqin, et al.
Pubblicazione: (2025)
Boundary-Aware Test-Time Adaptation for Zero-Shot Medical Image Segmentation
di: Xu, Chenlin, et al.
Pubblicazione: (2025)
di: Xu, Chenlin, et al.
Pubblicazione: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Adapter-Enhanced Semantic Prompting for Continual Learning
di: Yin, Baocai, et al.
Pubblicazione: (2024)
di: Yin, Baocai, et al.
Pubblicazione: (2024)
SPECTRUM: Semantic Processing and Emotion-informed video-Captioning Through Retrieval and Understanding Modalities
di: Faghihi, Ehsan, et al.
Pubblicazione: (2024)
di: Faghihi, Ehsan, et al.
Pubblicazione: (2024)
Retrieval-Augmented Egocentric Video Captioning
di: Xu, Jilan, et al.
Pubblicazione: (2024)
di: Xu, Jilan, et al.
Pubblicazione: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
di: Zhao, Shuai, et al.
Pubblicazione: (2023)
MatRes: Zero-Shot Test-Time Model Adaptation for Simultaneous Matching and Restoration
di: Lee, Kanggeon, et al.
Pubblicazione: (2026)
di: Lee, Kanggeon, et al.
Pubblicazione: (2026)
BaFTA: Backprop-Free Test-Time Adaptation For Zero-Shot Vision-Language Models
di: Hu, Xuefeng, et al.
Pubblicazione: (2024)
di: Hu, Xuefeng, et al.
Pubblicazione: (2024)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
di: Lebailly, Tim, et al.
Pubblicazione: (2025)
di: Lebailly, Tim, et al.
Pubblicazione: (2025)
Collaborative Temporal Consistency Learning for Point-supervised Natural Language Video Localization
di: Tao, Zhuo, et al.
Pubblicazione: (2025)
di: Tao, Zhuo, et al.
Pubblicazione: (2025)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
di: Jiang, Huajie, et al.
Pubblicazione: (2025)
di: Jiang, Huajie, et al.
Pubblicazione: (2025)
One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
di: Bianchi, Lorenzo, et al.
Pubblicazione: (2025)
di: Bianchi, Lorenzo, et al.
Pubblicazione: (2025)
Unlocking Prototype Potential: An Efficient Tuning Framework for Few-Shot Class-Incremental Learning
di: Jiang, Shengqin, et al.
Pubblicazione: (2026)
di: Jiang, Shengqin, et al.
Pubblicazione: (2026)
Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation
di: Zhu, Jingmin, et al.
Pubblicazione: (2025)
di: Zhu, Jingmin, et al.
Pubblicazione: (2025)
A Simple yet Effective Test-Time Adaptation for Zero-Shot Monocular Metric Depth Estimation
di: Marsal, Rémi, et al.
Pubblicazione: (2024)
di: Marsal, Rémi, et al.
Pubblicazione: (2024)
Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
di: Zeng, Runhao, et al.
Pubblicazione: (2025)
di: Zeng, Runhao, et al.
Pubblicazione: (2025)
A Comprehensive Survey of 3D Dense Captioning: Localizing and Describing Objects in 3D Scenes
di: Yu, Ting, et al.
Pubblicazione: (2024)
di: Yu, Ting, et al.
Pubblicazione: (2024)
Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images
di: Lu, Zimao, et al.
Pubblicazione: (2025)
di: Lu, Zimao, et al.
Pubblicazione: (2025)
Test-Time Zero-Shot Temporal Action Localization
di: Liberatori, Benedetta, et al.
Pubblicazione: (2024)
di: Liberatori, Benedetta, et al.
Pubblicazione: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
di: Yu, Eric Yang, et al.
Pubblicazione: (2024)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
di: Ding, Ning, et al.
Pubblicazione: (2025)
di: Ding, Ning, et al.
Pubblicazione: (2025)
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
di: Imam, Raza, et al.
Pubblicazione: (2024)
di: Imam, Raza, et al.
Pubblicazione: (2024)
DreamInsert: Zero-Shot Image-to-Video Object Insertion from A Single Image
di: Zhao, Qi, et al.
Pubblicazione: (2025)
di: Zhao, Qi, et al.
Pubblicazione: (2025)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
di: Liu, Mingyu, et al.
Pubblicazione: (2026)
di: Liu, Mingyu, et al.
Pubblicazione: (2026)
Uncertainty-boosted Robust Video Activity Anticipation
di: Qi, Zhaobo, et al.
Pubblicazione: (2024)
di: Qi, Zhaobo, et al.
Pubblicazione: (2024)
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
di: Jeon, MinJu, et al.
Pubblicazione: (2025)
Not All Pairs are Equal: Hierarchical Learning for Average-Precision-Oriented Video Retrieval
di: Liu, Yang, et al.
Pubblicazione: (2024)
di: Liu, Yang, et al.
Pubblicazione: (2024)
Zo3T: Zero-Shot 3D-Aware Trajectory-Guided Image-to-Video Generation via Test-Time Training
di: Zhang, Ruicheng, et al.
Pubblicazione: (2025)
di: Zhang, Ruicheng, et al.
Pubblicazione: (2025)
TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
di: Yao, Linli, et al.
Pubblicazione: (2026)
di: Yao, Linli, et al.
Pubblicazione: (2026)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
di: Tu, Yunbin, et al.
Pubblicazione: (2024)
Beyond Caption-Based Queries for Video Moment Retrieval
di: Pujol-Perich, David, et al.
Pubblicazione: (2026)
di: Pujol-Perich, David, et al.
Pubblicazione: (2026)
Documenti analoghi
-
SOVC: Subject-Oriented Video Captioning
di: Teng, Chang, et al.
Pubblicazione: (2023) -
Exploring the Temporal Consistency for Point-Level Weakly-Supervised Temporal Action Localization
di: Ma, Yunchuan, et al.
Pubblicazione: (2026) -
Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
di: Ma, Yunchuan, et al.
Pubblicazione: (2026) -
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
di: Tian, Mingkai, et al.
Pubblicazione: (2025) -
SDVPT: Semantic-Driven Visual Prompt Tuning for Open-World Object Counting
di: Zhao, Yiming, et al.
Pubblicazione: (2025)