Saved in:
| Main Authors: | Li, Senmao, Wang, Kai, Khan, Salman, Khan, Fahad Shahbaz, Yang, Jian, Wang, Yaxing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.16483 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diversity Has Always Been There in Your Visual Autoregressive Models
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024)
by: Li, Senmao, et al.
Published: (2024)
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
by: Li, Senmao, et al.
Published: (2025)
by: Li, Senmao, et al.
Published: (2025)
StyleDiffusion: Prompt-Embedding Inversion for Text-Based Editing
by: Li, Senmao, et al.
Published: (2023)
by: Li, Senmao, et al.
Published: (2023)
Towards Evaluating the Robustness of Visual State Space Models
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Vocabulary-free Fine-grained Visual Recognition via Enriched Contextually Grounded Vision-Language Model
by: Demidov, Dmitry, et al.
Published: (2025)
by: Demidov, Dmitry, et al.
Published: (2025)
Interpretable Zero-Shot Learning with Locally-Aligned Vision-Language Model
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
GroupMamba: Efficient Group-Based Visual State Space Model
by: Shaker, Abdelrahman, et al.
Published: (2024)
by: Shaker, Abdelrahman, et al.
Published: (2024)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
by: Maaz, Muhammad, et al.
Published: (2023)
by: Maaz, Muhammad, et al.
Published: (2023)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
Language Guided Domain Generalized Medical Image Segmentation
by: Kunhimon, Shahina, et al.
Published: (2024)
by: Kunhimon, Shahina, et al.
Published: (2024)
Cross-Modal Self-Training: Aligning Images and Pointclouds to Learn Classification without Labels
by: Dharmasiri, Amaya, et al.
Published: (2024)
by: Dharmasiri, Amaya, et al.
Published: (2024)
WorldCache: Content-Aware Caching for Accelerated Video World Models
by: Nawaz, Umair, et al.
Published: (2026)
by: Nawaz, Umair, et al.
Published: (2026)
Enhancing Novel Object Detection via Cooperative Foundational Models
by: Bharadwaj, Rohit, et al.
Published: (2023)
by: Bharadwaj, Rohit, et al.
Published: (2023)
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
by: Chen, Shiming, et al.
Published: (2025)
by: Chen, Shiming, et al.
Published: (2025)
Progressive Semantic-Guided Vision Transformer for Zero-Shot Learning
by: Chen, Shiming, et al.
Published: (2024)
by: Chen, Shiming, et al.
Published: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
by: Munasinghe, Shehan, et al.
Published: (2024)
by: Munasinghe, Shehan, et al.
Published: (2024)
Hierarchical Self-Supervised Adversarial Training for Robust Vision Models in Histopathology
by: Malik, Hashmat Shadab, et al.
Published: (2025)
by: Malik, Hashmat Shadab, et al.
Published: (2025)
VURF: A General-purpose Reasoning and Self-refinement Framework for Video Understanding
by: Mahmood, Ahmad, et al.
Published: (2024)
by: Mahmood, Ahmad, et al.
Published: (2024)
VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
by: Bharadwaj, Rohit, et al.
Published: (2024)
by: Bharadwaj, Rohit, et al.
Published: (2024)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
TAViS: Text-bridged Audio-Visual Segmentation with Foundation Models
by: Luo, Ziyang, et al.
Published: (2025)
by: Luo, Ziyang, et al.
Published: (2025)
Mobile-VideoGPT: Fast and Accurate Model for Mobile Video Understanding
by: Shaker, Abdelrahman, et al.
Published: (2025)
by: Shaker, Abdelrahman, et al.
Published: (2025)
Concept Drift and Long-Tailed Distribution in Fine-Grained Visual Categorization: Benchmark and Method
by: Ye, Shuo, et al.
Published: (2023)
by: Ye, Shuo, et al.
Published: (2023)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
by: Wasim, Syed Talal, et al.
Published: (2023)
by: Wasim, Syed Talal, et al.
Published: (2023)
RainDiff: End-to-end Precipitation Nowcasting Via Token-wise Attention Diffusion
by: Nguyen, Thao, et al.
Published: (2025)
by: Nguyen, Thao, et al.
Published: (2025)
UniMed-CLIP: Towards a Unified Image-Text Pretraining Paradigm for Diverse Medical Imaging Modalities
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
by: Kumar, Komal, et al.
Published: (2025)
by: Kumar, Komal, et al.
Published: (2025)
WaDi: Weight Direction-aware Distillation for One-step Image Synthesis
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device
by: Shaker, Abdelrahman, et al.
Published: (2026)
by: Shaker, Abdelrahman, et al.
Published: (2026)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
by: Shaker, Abdelrahman, et al.
Published: (2022)
by: Shaker, Abdelrahman, et al.
Published: (2022)
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
by: Dong, Jiahua, et al.
Published: (2025)
by: Dong, Jiahua, et al.
Published: (2025)
Visual-Augmented Dynamic Semantic Prototype for Generative Zero-Shot Learning
by: Hou, Wenjin, et al.
Published: (2024)
by: Hou, Wenjin, et al.
Published: (2024)
Hierarchical Text-to-Vision Self Supervised Alignment for Improved Histopathology Representation Learning
by: Watawana, Hasindri, et al.
Published: (2024)
by: Watawana, Hasindri, et al.
Published: (2024)
Similar Items
-
Diversity Has Always Been There in Your Visual Autoregressive Models
by: Wang, Tong, et al.
Published: (2025) -
Faster Diffusion: Rethinking the Role of the Encoder for Diffusion Model Inference
by: Li, Senmao, et al.
Published: (2023) -
Get What You Want, Not What You Don't: Image Content Suppression for Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2024) -
One-Way Ticket:Time-Independent Unified Encoder for Distilling Text-to-Image Diffusion Models
by: Li, Senmao, et al.
Published: (2025) -
InterLCM: Low-Quality Images as Intermediate States of Latent Consistency Models for Effective Blind Face Restoration
by: Li, Senmao, et al.
Published: (2025)