Bring Your Dreams to Life: Continual Text-to-Video Customization
Fuente:
arXiv
Guardado en:
| Autores principales: | Dong, Jiahua, Wang, Xudong, Liang, Wenqi, Han, Zongyan, Cao, Meng, Zhang, Duzhen, Zhao, Hanbin, Han, Zhi, Khan, Salman, Khan, Fahad Shahbaz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
por: Dong, Jiahua, et al.
Publicado: (2024)
por: Dong, Jiahua, et al.
Publicado: (2024)
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
por: Dong, Jiahua, et al.
Publicado: (2025)
por: Dong, Jiahua, et al.
Publicado: (2025)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
por: Wang, Tong, et al.
Publicado: (2025)
por: Wang, Tong, et al.
Publicado: (2025)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
por: Munasinghe, Shehan, et al.
Publicado: (2024)
por: Munasinghe, Shehan, et al.
Publicado: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
por: Maaz, Muhammad, et al.
Publicado: (2023)
por: Maaz, Muhammad, et al.
Publicado: (2023)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
por: Maaz, Muhammad, et al.
Publicado: (2025)
por: Maaz, Muhammad, et al.
Publicado: (2025)
Connecting Dreams with Visual Brainstorming Instruction
por: Sun, Yasheng, et al.
Publicado: (2024)
por: Sun, Yasheng, et al.
Publicado: (2024)
Efficient Localized Adaptation of Neural Weather Forecasting: A Case Study in the MENA Region
por: Munir, Muhammad Akhtar, et al.
Publicado: (2024)
por: Munir, Muhammad Akhtar, et al.
Publicado: (2024)
Dual Hyperspectral Mamba for Efficient Spectral Compressive Imaging
por: Dong, Jiahua, et al.
Publicado: (2024)
por: Dong, Jiahua, et al.
Publicado: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
por: Bharadwaj, Rohit, et al.
Publicado: (2024)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
por: Mahmood, Ahmad, et al.
Publicado: (2024)
por: Mahmood, Ahmad, et al.
Publicado: (2024)
WorldCache: Content-Aware Caching for Accelerated Video World Models
por: Nawaz, Umair, et al.
Publicado: (2026)
por: Nawaz, Umair, et al.
Publicado: (2026)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
por: Shaker, Abdelrahman, et al.
Publicado: (2025)
por: Shaker, Abdelrahman, et al.
Publicado: (2025)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
por: Wasim, Syed Talal, et al.
Publicado: (2023)
por: Wasim, Syed Talal, et al.
Publicado: (2023)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
Language Guided Domain Generalized Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2024)
por: Kunhimon, Shahina, et al.
Publicado: (2024)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
por: Chen, Shiming, et al.
Publicado: (2025)
por: Chen, Shiming, et al.
Publicado: (2025)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
por: Dharmasiri, Amaya, et al.
Publicado: (2024)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
por: Chen, Shiming, et al.
Publicado: (2024)
por: Chen, Shiming, et al.
Publicado: (2024)
Enhancing Novel Object Detection via Cooperative Foundational Models
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
por: Bharadwaj, Rohit, et al.
Publicado: (2023)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
por: Rasheed, Hanoona, et al.
Publicado: (2025)
por: Rasheed, Hanoona, et al.
Publicado: (2025)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
por: Watawana, Hasindri, et al.
Publicado: (2024)
por: Watawana, Hasindri, et al.
Publicado: (2024)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
por: Maaz, Muhammad, et al.
Publicado: (2024)
por: Maaz, Muhammad, et al.
Publicado: (2024)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
por: Kumar, Komal, et al.
Publicado: (2025)
por: Kumar, Komal, et al.
Publicado: (2025)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
por: Luo, Ziyang, et al.
Publicado: (2025)
por: Luo, Ziyang, et al.
Publicado: (2025)
Align Your Prompts: Test-Time Prompting with Distribution Alignment for Zero-Shot Generalization
por: Hassan, Jameel, et al.
Publicado: (2023)
por: Hassan, Jameel, et al.
Publicado: (2023)
Modulate Your Spectrum in Self-Supervised Learning
por: Weng, Xi, et al.
Publicado: (2023)
por: Weng, Xi, et al.
Publicado: (2023)
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
por: Liu, Baichen, et al.
Publicado: (2025)
por: Liu, Baichen, et al.
Publicado: (2025)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
por: Demidov, Dmitry, et al.
Publicado: (2025)
por: Demidov, Dmitry, et al.
Publicado: (2025)
Learnable Weight Initialization for Volumetric Medical Image Segmentation
por: Kunhimon, Shahina, et al.
Publicado: (2023)
por: Kunhimon, Shahina, et al.
Publicado: (2023)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
por: Noman, Mubashir, et al.
Publicado: (2024)
por: Noman, Mubashir, et al.
Publicado: (2024)
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
por: Kumar, Komal, et al.
Publicado: (2026)
por: Kumar, Komal, et al.
Publicado: (2026)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
por: Rasheed, Hanoona, et al.
Publicado: (2025)
por: Rasheed, Hanoona, et al.
Publicado: (2025)
Efficient Video Object Segmentation via Modulated Cross-Attention Memory
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
por: Shaker, Abdelrahman, et al.
Publicado: (2024)
GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support
por: Sheikh, Muhammad Umer, et al.
Publicado: (2026)
por: Sheikh, Muhammad Umer, et al.
Publicado: (2026)
Towards Evaluating the Robustness of Visual State Space Models
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
por: Malik, Hashmat Shadab, et al.
Publicado: (2024)
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
por: Khattak, Muhammad Uzair, et al.
Publicado: (2024)
VideoMolmo: Spatio-Temporal Grounding Meets Pointing
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
por: Ahmad, Ghazi Shazan, et al.
Publicado: (2025)
MuseumMaker: Continual Style Customization without Catastrophic Forgetting
por: Liu, Chenxi, et al.
Publicado: (2024)
por: Liu, Chenxi, et al.
Publicado: (2024)
Ejemplares similares
-
How to Continually Adapt Text-to-Image Diffusion Models for Flexible Customization?
por: Dong, Jiahua, et al.
Publicado: (2024) -
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
por: Dong, Jiahua, et al.
Publicado: (2025) -
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
por: Wang, Tong, et al.
Publicado: (2025) -
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
por: Munasinghe, Shehan, et al.
Publicado: (2024) -
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
por: Maaz, Muhammad, et al.
Publicado: (2023)