Frame-Difference Guided Dynamic Region Perception for CLIP Adaptation in Text-Video Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yu, Jiaao, Han, Mingjie, Gong, Tao, Zhang, Jian, Lan, Man |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
von: Yu, Jiaao, et al.
Veröffentlicht: (2025)
UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
von: Liu, Delong, et al.
Veröffentlicht: (2024)
von: Liu, Delong, et al.
Veröffentlicht: (2024)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024)
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
von: Wang, Ni, et al.
Veröffentlicht: (2024)
von: Wang, Ni, et al.
Veröffentlicht: (2024)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
von: Pan, Jiancheng, et al.
Veröffentlicht: (2024)
von: Pan, Jiancheng, et al.
Veröffentlicht: (2024)
Focus, Distinguish, and Prompt: Unleashing CLIP for Efficient and Flexible Scene Text Retrieval
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
von: Zeng, Gangyan, et al.
Veröffentlicht: (2024)
CLIP-based Synergistic Knowledge Transfer for Text-based Person Retrieval
von: Liu, Yating, et al.
Veröffentlicht: (2023)
von: Liu, Yating, et al.
Veröffentlicht: (2023)
MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval
von: Elallaf, Ahmad, et al.
Veröffentlicht: (2026)
von: Elallaf, Ahmad, et al.
Veröffentlicht: (2026)
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
von: Ye, Bo, et al.
Veröffentlicht: (2026)
von: Ye, Bo, et al.
Veröffentlicht: (2026)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
von: Yang, Bowen, et al.
Veröffentlicht: (2025)
Explicit Uncertainty Modeling for Active CLIP Adaptation with Dual Prompt Tuning
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
Investigation of Frame Differences as Motion Cues for Video Object Segmentation
von: Kawamura, Sota, et al.
Veröffentlicht: (2025)
von: Kawamura, Sota, et al.
Veröffentlicht: (2025)
CLIP-Guided Unsupervised Semantic-Aware Exposure Correction
von: Wu, Puzhen, et al.
Veröffentlicht: (2026)
von: Wu, Puzhen, et al.
Veröffentlicht: (2026)
Unleash the Potential of CLIP for Video Highlight Detection
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
von: Han, Donghoon, et al.
Veröffentlicht: (2024)
DualCross: Cross-Modality Cross-Domain Adaptation for Monocular BEV Perception
von: Man, Yunze, et al.
Veröffentlicht: (2023)
von: Man, Yunze, et al.
Veröffentlicht: (2023)
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
von: Zhang, Zhiqiu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiqiu, et al.
Veröffentlicht: (2025)
Text-Guided Multi-Scale Frequency Representation Adaptation
von: Yan, Weicai, et al.
Veröffentlicht: (2026)
von: Yan, Weicai, et al.
Veröffentlicht: (2026)
DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihang, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Understanding with CLIP-Based Image-to-Text Transformation
von: Che, Chang, et al.
Veröffentlicht: (2024)
von: Che, Chang, et al.
Veröffentlicht: (2024)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
von: Tzachor, Issar, et al.
Veröffentlicht: (2026)
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
von: Zhang, Peng, et al.
Veröffentlicht: (2026)
von: Zhang, Peng, et al.
Veröffentlicht: (2026)
ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
von: Tran, Quoc-Khang, et al.
Veröffentlicht: (2026)
A Denoising Framework for Real-World Ultra-Low-Dose Lung CT Images Based on an Image Purification Strategy
von: Gong, Guoliang, et al.
Veröffentlicht: (2025)
von: Gong, Guoliang, et al.
Veröffentlicht: (2025)
IPv2: An Improved Image Purification Strategy for Real-World Ultra-Low-Dose Lung CT Denoising
von: Gong, Guoliang, et al.
Veröffentlicht: (2026)
von: Gong, Guoliang, et al.
Veröffentlicht: (2026)
HarmoCLIP: Harmonizing Global and Regional Representations in Contrastive Vision-Language Models
von: Zeng, Haoxi, et al.
Veröffentlicht: (2025)
von: Zeng, Haoxi, et al.
Veröffentlicht: (2025)
EgoPrune: Efficient Token Pruning for Egomotion Video Reasoning in Embodied Agent
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
von: Li, Jiaao, et al.
Veröffentlicht: (2025)
Parrot Captions Teach CLIP to Spot Text
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
von: Lin, Yiqi, et al.
Veröffentlicht: (2023)
CLIP-MUSED: CLIP-Guided Multi-Subject Visual Neural Information Semantic Decoding
von: Zhou, Qiongyi, et al.
Veröffentlicht: (2024)
von: Zhou, Qiongyi, et al.
Veröffentlicht: (2024)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
von: Kim, Junho, et al.
Veröffentlicht: (2024)
von: Kim, Junho, et al.
Veröffentlicht: (2024)
Beyond Pedestrians: Caption-Guided CLIP Framework for High-Difficulty Video-based Person Re-Identification
von: Hamano, Shogo, et al.
Veröffentlicht: (2026)
von: Hamano, Shogo, et al.
Veröffentlicht: (2026)
Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval
von: Cho, CH, et al.
Veröffentlicht: (2025)
von: Cho, CH, et al.
Veröffentlicht: (2025)
microCLIP: Unsupervised CLIP Adaptation via Coarse-Fine Token Fusion for Fine-Grained Image Classification
von: Silva, Sathira, et al.
Veröffentlicht: (2025)
von: Silva, Sathira, et al.
Veröffentlicht: (2025)
UrbanCross: Enhancing Satellite Image-Text Retrieval with Cross-Domain Adaptation
von: Zhong, Siru, et al.
Veröffentlicht: (2024)
von: Zhong, Siru, et al.
Veröffentlicht: (2024)
Dynamic Eraser for Guided Concept Erasure in Diffusion Models
von: Gong, Qinghui
Veröffentlicht: (2026)
von: Gong, Qinghui
Veröffentlicht: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
Detecting AI-Generated Video via Frame Consistency
von: Ma, Long, et al.
Veröffentlicht: (2024)
von: Ma, Long, et al.
Veröffentlicht: (2024)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
von: Wang, Zhengbo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Semantic Relation-Enhanced CLIP Adapter for Domain Adaptive Zero-Shot Learning
von: Yu, Jiaao, et al.
Veröffentlicht: (2025) -
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
von: Yu, Jiaao, et al.
Veröffentlicht: (2025) -
UFO: Enhancing Diffusion-Based Video Generation with a Uniform Frame Organizer
von: Liu, Delong, et al.
Veröffentlicht: (2024) -
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
von: Jing, Xiaolun, et al.
Veröffentlicht: (2024) -
Multi-Scale Temporal Difference Transformer for Video-Text Retrieval
von: Wang, Ni, et al.
Veröffentlicht: (2024)