Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Dong, Jiahua, Yin, Hui, Liang, Wenqi, Zhao, Hanbin, Ding, Henghui, Sebe, Nicu, Khan, Salman, Khan, Fahad Shahbaz |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
par: Dong, Jiahua, et autres
Publié: (2024)
par: Dong, Jiahua, et autres
Publié: (2024)
IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting
par: Fu, Hao, et autres
Publié: (2025)
par: Fu, Hao, et autres
Publié: (2025)
Bring Your Dreams to Life: Continual Text-to-Video Customization
par: Dong, Jiahua, et autres
Publié: (2025)
par: Dong, Jiahua, et autres
Publié: (2025)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
par: Munasinghe, Shehan, et autres
Publié: (2024)
par: Munasinghe, Shehan, et autres
Publié: (2024)
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
par: Dong, Jiahua, et autres
Publié: (2024)
par: Dong, Jiahua, et autres
Publié: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
par: Maaz, Muhammad, et autres
Publié: (2023)
par: Maaz, Muhammad, et autres
Publié: (2023)
Language Guided Domain Generalized Medical Image Segmentation
par: Kunhimon, Shahina, et autres
Publié: (2024)
par: Kunhimon, Shahina, et autres
Publié: (2024)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
par: Maaz, Muhammad, et autres
Publié: (2025)
par: Maaz, Muhammad, et autres
Publié: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
par: Chen, Shiming, et autres
Publié: (2025)
par: Chen, Shiming, et autres
Publié: (2025)
PECTP: Parameter-Efficient Cross-Task Prompts for Incremental Vision Transformer
par: Feng, Qian, et autres
Publié: (2024)
par: Feng, Qian, et autres
Publié: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
par: Dharmasiri, Amaya, et autres
Publié: (2024)
par: Dharmasiri, Amaya, et autres
Publié: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
par: Watawana, Hasindri, et autres
Publié: (2024)
par: Watawana, Hasindri, et autres
Publié: (2024)
CE-SDWV: Effective and Efficient Concept Erasure for Text-to-Image Diffusion Models via a Semantic-Driven Word Vocabulary
par: Tu, Jiahang, et autres
Publié: (2025)
par: Tu, Jiahang, et autres
Publié: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
par: Chen, Shiming, et autres
Publié: (2024)
par: Chen, Shiming, et autres
Publié: (2024)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
par: Chen, Shiming, et autres
Publié: (2025)
par: Chen, Shiming, et autres
Publié: (2025)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
par: Bharadwaj, Rohit, et autres
Publié: (2024)
par: Bharadwaj, Rohit, et autres
Publié: (2024)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
par: Shaker, Abdelrahman, et autres
Publié: (2024)
par: Shaker, Abdelrahman, et autres
Publié: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
par: Malik, Hashmat Shadab, et autres
Publié: (2025)
par: Malik, Hashmat Shadab, et autres
Publié: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
par: Kunhimon, Shahina, et autres
Publié: (2023)
par: Kunhimon, Shahina, et autres
Publié: (2023)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
par: Mahmood, Ahmad, et autres
Publié: (2024)
par: Mahmood, Ahmad, et autres
Publié: (2024)
Towards Evaluating the Robustness of Visual State Space Models
par: Malik, Hashmat Shadab, et autres
Publié: (2024)
par: Malik, Hashmat Shadab, et autres
Publié: (2024)
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
par: Boudjoghra, Mohamed El Amine, et autres
Publié: (2024)
par: Boudjoghra, Mohamed El Amine, et autres
Publié: (2024)
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
par: Liu, Baichen, et autres
Publié: (2025)
par: Liu, Baichen, et autres
Publié: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
par: Demidov, Dmitry, et autres
Publié: (2025)
par: Demidov, Dmitry, et autres
Publié: (2025)
BAPLe: Backdoor Attacks on Medical Foundational Models using Prompt Learning
par: Hanif, Asif, et autres
Publié: (2024)
par: Hanif, Asif, et autres
Publié: (2024)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
par: Shaker, Abdelrahman, et autres
Publié: (2025)
par: Shaker, Abdelrahman, et autres
Publié: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
par: Wasim, Syed Talal, et autres
Publié: (2023)
par: Wasim, Syed Talal, et autres
Publié: (2023)
GroupMamba: Efficient Group-Based Visual State Space Model
par: Shaker, Abdelrahman, et autres
Publié: (2024)
par: Shaker, Abdelrahman, et autres
Publié: (2024)
AURORA:Augmented Understanding via Structured Reasoning and Reinforcement Learning for Reference Audio-Visual Segmentation
par: Luo, Ziyang, et autres
Publié: (2025)
par: Luo, Ziyang, et autres
Publié: (2025)
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
par: He, Shuting, et autres
Publié: (2024)
par: He, Shuting, et autres
Publié: (2024)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
par: Rasheed, Hanoona, et autres
Publié: (2025)
par: Rasheed, Hanoona, et autres
Publié: (2025)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
par: Luo, Ziyang, et autres
Publié: (2025)
par: Luo, Ziyang, et autres
Publié: (2025)
WorldCache: Content-Aware Caching for Accelerated Video World Models
par: Nawaz, Umair, et autres
Publié: (2026)
par: Nawaz, Umair, et autres
Publié: (2026)
FasterVAR: Plug-and-Play Acceleration for Visual Autoregressive Models
par: Li, Senmao, et autres
Publié: (2025)
par: Li, Senmao, et autres
Publié: (2025)
Enhancing Novel Object Detection via Cooperative Foundational Models
par: Bharadwaj, Rohit, et autres
Publié: (2023)
par: Bharadwaj, Rohit, et autres
Publié: (2023)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
par: Maaz, Muhammad, et autres
Publié: (2024)
par: Maaz, Muhammad, et autres
Publié: (2024)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
par: Hassan, Jameel, et autres
Publié: (2023)
par: Hassan, Jameel, et autres
Publié: (2023)
MedContext: Learning Contextual Cues for Efficient Volumetric Medical Segmentation
par: Gani, Hanan, et autres
Publié: (2024)
par: Gani, Hanan, et autres
Publié: (2024)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
par: Shaker, Abdelrahman, et autres
Publié: (2022)
par: Shaker, Abdelrahman, et autres
Publié: (2022)
LW2G: Learning Whether to Grow for Prompt-based Continual Learning
par: Feng, Qian, et autres
Publié: (2024)
par: Feng, Qian, et autres
Publié: (2024)
Documents similaires
-
How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
par: Dong, Jiahua, et autres
Publié: (2024) -
IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting
par: Fu, Hao, et autres
Publié: (2025) -
Bring Your Dreams to Life: Continual Text-to-Video Customization
par: Dong, Jiahua, et autres
Publié: (2025) -
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
par: Munasinghe, Shehan, et autres
Publié: (2024) -
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
par: Dong, Jiahua, et autres
Publié: (2024)